You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Pyglet获取media.load()加载音频的频率以实现可视化?

Yes, You Can Extract Audio Frequencies with Pyglet for Visualization!

Absolutely! Pyglet fully supports pulling audio frequency data from files loaded via media.load()—you just need to combine it with some basic audio processing (using NumPy for FFT, since Pyglet doesn’t include built-in frequency analysis tools). Here’s how to make it work:

Step 1: Load Audio and Extract Raw Samples

First, load your audio file and grab the raw sample data. Pyglet’s media.load() returns an audio source that lets you access buffers of sample data. Note: WAV files work seamlessly, but MP3s may require FFmpeg integration (ensure your Pyglet setup has FFmpeg support for compressed formats).

import pyglet
import numpy as np

# Load your audio file
audio_source = pyglet.media.load("your_audio.wav")
audio_format = audio_source.audio_format  # Get sample rate, channels, bit depth

# Extract raw audio samples from the source's buffers
raw_samples = []
for buffer in audio_source.get_buffers():
    raw_samples.extend(buffer.data)

# Convert samples to a NumPy array (adjust dtype based on your audio's bit depth)
# 16-bit PCM is the most common, so we'll use np.int16 here
sample_array = np.array(raw_samples, dtype=np.int16)

# Handle stereo audio (split into channels if needed; use one for analysis)
if audio_format.channels == 2:
    left_channel = sample_array[::2]
    right_channel = sample_array[1::2]
    analysis_data = left_channel  # Use left channel for frequency analysis
else:
    analysis_data = sample_array

Step 2: Calculate Frequency Data with FFT

To get frequency values, you’ll use Fast Fourier Transform (FFT) to convert the time-domain audio samples into frequency-domain data. NumPy’s fft module makes this straightforward. We’ll process the audio in small frames (for smooth visualization) and compute magnitude values (converted to decibels for better dynamic range).

# Define frame and hop sizes (tweak these for your visualization speed/resolution)
frame_size = 1024  # Larger frames = more frequency detail, slower updates
hop_length = 512   # Distance between start of each frame

# Get frequency bins (matches the FFT output)
frequency_bins = np.fft.fftfreq(frame_size, d=1/audio_format.sample_rate)
positive_bins = frequency_bins[:frame_size//2]  # Ignore negative frequencies (mirror of positive)

# Process each frame to generate frequency data
for i in range(0, len(analysis_data) - frame_size, hop_length):
    # Extract a single frame of audio
    frame = analysis_data[i:i+frame_size]
    
    # Apply a Hanning window to reduce spectral leakage
    windowed_frame = frame * np.hanning(frame_size)
    
    # Compute FFT and magnitude spectrum
    fft_result = np.fft.fft(windowed_frame)
    magnitude = np.abs(fft_result[:frame_size//2])
    
    # Convert magnitude to decibels (avoids log(0) with a small offset)
    magnitude_db = 20 * np.log10(magnitude + 1e-10)
    
    # Now you can use `positive_bins` (frequency values) and `magnitude_db` (amplitude) for visualization!

Step 3: Build a Real-Time Visualization (Optional)

If you want to visualize frequencies as the audio plays, integrate the frequency calculation into Pyglet’s event loop. Here’s a quick example of drawing bar graphs that update with the audio:

# Set up a Pyglet window for visualization
window = pyglet.window.Window(width=800, height=400)
batch = pyglet.graphics.Batch()
bars = []

# Initialize bar graphics (reduce bin count for better visibility)
num_bars = len(positive_bins) // 10
bar_width = window.width / num_bars

for i in range(num_bars):
    x_pos = i * bar_width
    # Create a quad for each bar (initial height = 0)
    bars.append(batch.add(4, pyglet.gl.GL_QUADS, None,
                          ('v2f', (x_pos, 0, x_pos+bar_width, 0, x_pos+bar_width, 10, x_pos, 10)),
                          ('c3f', (0.5, 0.8, 1.0)*4)))  # Light blue color

def update_bar_heights(magnitude_db):
    # Scale decibel values to fit the window height
    max_db = np.max(magnitude_db)
    min_db = np.min(magnitude_db)
    scaled_heights = (magnitude_db - min_db) / (max_db - min_db) * window.height
    
    # Update each bar's height
    for i in range(num_bars):
        bin_index = i * (len(magnitude_db) // num_bars)
        height = scaled_heights[bin_index] if bin_index < len(scaled_heights) else 0
        bar = bars[i]
        bar.vertices = (i*bar_width, 0, (i+1)*bar_width, 0, (i+1)*bar_width, height, i*bar_width, height)

@window.event
def on_draw():
    window.clear()
    batch.draw()

# Track our position in the audio sample array
current_frame_idx = 0

def process_audio_frame(dt):
    global current_frame_idx
    if current_frame_idx + frame_size < len(analysis_data):
        # Calculate frequency data for the current frame
        frame = analysis_data[current_frame_idx:current_frame_idx+frame_size]
        windowed_frame = frame * np.hanning(frame_size)
        fft_result = np.fft.fft(windowed_frame)
        magnitude_db = 20 * np.log10(np.abs(fft_result[:frame_size//2]) + 1e-10)
        
        # Update the visualization
        update_bar_heights(magnitude_db)
        
        # Move to the next frame
        current_frame_idx += hop_length

# Schedule updates at 30 FPS (matches typical visualization smoothness)
pyglet.clock.schedule_interval(process_audio_frame, 1/30)

# Run the Pyglet app
pyglet.app.run()

Key Notes to Keep in Mind

  • Audio Decoding: For compressed formats like MP3, ensure Pyglet is configured to use FFmpeg (you may need to install pyglet[ffmpeg] via pip).
  • Performance: Larger frame_size values give more frequency detail but slower updates. Adjust based on your hardware and desired visualization speed.
  • Real-Time Playback: To sync visualization with audio playback, use pyglet.media.Player and capture frames as the audio plays—you can hook into the player’s buffer events or periodically sample its current state.

内容的提问来源于stack exchange,提问作者Rael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:24:05