Librosa绘制音频频谱图显示占比差异问题求助——采样率调整无改善
Let's break down your problem first: You've got two audio datasets—one at 22kHz with a spectrogram where the purple (high-energy) region takes up ~1/3 of the image, and another at 44.1kHz where purple fills nearly the whole frame. But when you downsample the 44.1kHz audio using librosa.load(filename, sr=22000), the spectrogram looks identical to the original. Here's how to diagnose and fix this:
1. Verify Downsampling Actually Worked
First, make sure the downsampling step isn't failing silently. Don't just trust the spectrogram—check the raw audio data:
- Print the
srvalue right after loading to confirm it's 22000, not still 44100. - Compare the length of
y(the audio time series) before and after downsampling. For the same audio duration, the downsampled version should have roughly half the number of samples (since 44100 / 22000 ≈ 2). - Use
librosa.get_duration(y=y, sr=sr)to check that the audio duration stays consistent—if it changes, that means the load/downsample step isn't behaving as expected.
2. Explicitly Set Spectrogram Display Parameters
Your current code uses default settings for librosa.display.specshow, which might not be picking up the downsampled sr correctly. Try explicitly passing the sampling rate to ensure the frequency axis scales properly:
librosa.display.specshow(S_db, sr=sr)
For a 22kHz audio, the Nyquist frequency is 11kHz, so the frequency axis should only go up to 11kHz (half the 44.1kHz original's 22.05kHz range). Explicitly setting sr ensures the plot reflects this.
3. Check the Audio's Intrinsic Frequency Content
The purple regions represent high-energy frequencies, so the two datasets might just have fundamentally different frequency distributions:
- Your 22kHz dataset's audio might only contain low-frequency content (e.g., concentrated below ~3.6kHz, which is 1/3 of 11kHz), while the 44.1kHz dataset has energy spanning from low all the way up to ~22kHz.
- Downsampling should filter out frequencies above 11kHz, but you can confirm this by calculating spectral metrics:
A lower centroid after downsampling means high frequencies were successfully removed.# Calculate spectral centroid (average frequency of energy) centroid = librosa.feature.spectral_centroid(y=y, sr=sr) print(f"Spectral centroid: {np.mean(centroid)} Hz")
4. Standardize the dB Scale Reference
librosa.amplitude_to_db uses the maximum amplitude of the current audio as the reference by default. If your two datasets have very different overall energy levels, this can skew how the purple regions appear. Try using a fixed reference to align the dynamic range:
# Use a fixed reference value (e.g., 1.0, or the max amplitude from your 22kHz dataset) S_db = librosa.amplitude_to_db(np.abs(D), ref=1.0)
This makes it easier to compare the relative energy distribution across both datasets.
5. Validate the Spectrogram Data Before Saving
Sometimes the issue isn't with the spectrogram itself, but how it's saved or displayed. Try these checks:
- Add
plt.show()beforeplt.savefig()to view the spectrogram directly in your environment—this avoids any artifacts from the save process. - Compare the numerical values of
S_dbbefore and after downsampling: print its shape, max/min dB values, or slice a portion of the array to see if the high-frequency bins (the upper part of the spectrogram) are muted after downsampling.
Modified Example Code with Validation Steps
Here's an adjusted version of your code with built-in checks to help diagnose the issue:
import librosa import librosa.display import matplotlib.pyplot as plt import numpy as np filename = "your_44khz_audio.wav" output_file = "downsampled_spectrogram.png" # Load with downsampling y, sr = librosa.load(filename, sr=22000) print(f"Loaded sampling rate: {sr}") print(f"Audio sample count: {len(y)} | Duration: {librosa.get_duration(y=y, sr=sr):.2f}s") # Compute STFT and convert to dB D = librosa.stft(y) S_db = librosa.amplitude_to_db(np.abs(D), ref=np.max) print(f"Spectrogram shape: {S_db.shape}") print(f"Max dB value: {np.max(S_db):.2f} | Min dB value: {np.min(S_db):.2f}") # Generate and display spectrogram plt.figure(figsize=(3, 3), dpi=100, frameon=False) librosa.display.specshow(S_db, sr=sr) # Explicit sr parameter plt.subplots_adjust(top=1, bottom=0, right=1, left=0, hspace=0, wspace=0) plt.margins(0, 0) plt.gca().xaxis.set_major_locator(plt.NullLocator()) plt.gca().yaxis.set_major_locator(plt.NullLocator()) plt.axis('off') plt.show() # View first to confirm plt.savefig(output_file, dpi=100, pad_inches=0)
Start with step 1—confirming downsampling works is the foundation. If that checks out, move to verifying the frequency content and display settings.
内容的提问来源于stack exchange,提问作者Mattia Campana

