You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Librosa绘制音频频谱图显示占比差异问题求助——采样率调整无改善

Troubleshooting & Solutions for Librosa Spectrogram Sampling Rate Discrepancies

Let's break down your problem first: You've got two audio datasets—one at 22kHz with a spectrogram where the purple (high-energy) region takes up ~1/3 of the image, and another at 44.1kHz where purple fills nearly the whole frame. But when you downsample the 44.1kHz audio using librosa.load(filename, sr=22000), the spectrogram looks identical to the original. Here's how to diagnose and fix this:

1. Verify Downsampling Actually Worked

First, make sure the downsampling step isn't failing silently. Don't just trust the spectrogram—check the raw audio data:

  • Print the sr value right after loading to confirm it's 22000, not still 44100.
  • Compare the length of y (the audio time series) before and after downsampling. For the same audio duration, the downsampled version should have roughly half the number of samples (since 44100 / 22000 ≈ 2).
  • Use librosa.get_duration(y=y, sr=sr) to check that the audio duration stays consistent—if it changes, that means the load/downsample step isn't behaving as expected.

2. Explicitly Set Spectrogram Display Parameters

Your current code uses default settings for librosa.display.specshow, which might not be picking up the downsampled sr correctly. Try explicitly passing the sampling rate to ensure the frequency axis scales properly:

librosa.display.specshow(S_db, sr=sr)

For a 22kHz audio, the Nyquist frequency is 11kHz, so the frequency axis should only go up to 11kHz (half the 44.1kHz original's 22.05kHz range). Explicitly setting sr ensures the plot reflects this.

3. Check the Audio's Intrinsic Frequency Content

The purple regions represent high-energy frequencies, so the two datasets might just have fundamentally different frequency distributions:

  • Your 22kHz dataset's audio might only contain low-frequency content (e.g., concentrated below ~3.6kHz, which is 1/3 of 11kHz), while the 44.1kHz dataset has energy spanning from low all the way up to ~22kHz.
  • Downsampling should filter out frequencies above 11kHz, but you can confirm this by calculating spectral metrics:
    # Calculate spectral centroid (average frequency of energy)
    centroid = librosa.feature.spectral_centroid(y=y, sr=sr)
    print(f"Spectral centroid: {np.mean(centroid)} Hz")
    
    A lower centroid after downsampling means high frequencies were successfully removed.

4. Standardize the dB Scale Reference

librosa.amplitude_to_db uses the maximum amplitude of the current audio as the reference by default. If your two datasets have very different overall energy levels, this can skew how the purple regions appear. Try using a fixed reference to align the dynamic range:

# Use a fixed reference value (e.g., 1.0, or the max amplitude from your 22kHz dataset)
S_db = librosa.amplitude_to_db(np.abs(D), ref=1.0)

This makes it easier to compare the relative energy distribution across both datasets.

5. Validate the Spectrogram Data Before Saving

Sometimes the issue isn't with the spectrogram itself, but how it's saved or displayed. Try these checks:

  • Add plt.show() before plt.savefig() to view the spectrogram directly in your environment—this avoids any artifacts from the save process.
  • Compare the numerical values of S_db before and after downsampling: print its shape, max/min dB values, or slice a portion of the array to see if the high-frequency bins (the upper part of the spectrogram) are muted after downsampling.

Modified Example Code with Validation Steps

Here's an adjusted version of your code with built-in checks to help diagnose the issue:

import librosa
import librosa.display
import matplotlib.pyplot as plt
import numpy as np

filename = "your_44khz_audio.wav"
output_file = "downsampled_spectrogram.png"

# Load with downsampling
y, sr = librosa.load(filename, sr=22000)
print(f"Loaded sampling rate: {sr}")
print(f"Audio sample count: {len(y)} | Duration: {librosa.get_duration(y=y, sr=sr):.2f}s")

# Compute STFT and convert to dB
D = librosa.stft(y)
S_db = librosa.amplitude_to_db(np.abs(D), ref=np.max)
print(f"Spectrogram shape: {S_db.shape}")
print(f"Max dB value: {np.max(S_db):.2f} | Min dB value: {np.min(S_db):.2f}")

# Generate and display spectrogram
plt.figure(figsize=(3, 3), dpi=100, frameon=False)
librosa.display.specshow(S_db, sr=sr)  # Explicit sr parameter
plt.subplots_adjust(top=1, bottom=0, right=1, left=0, hspace=0, wspace=0)
plt.margins(0, 0)
plt.gca().xaxis.set_major_locator(plt.NullLocator())
plt.gca().yaxis.set_major_locator(plt.NullLocator())
plt.axis('off')

plt.show()  # View first to confirm
plt.savefig(output_file, dpi=100, pad_inches=0)

Start with step 1—confirming downsampling works is the foundation. If that checks out, move to verifying the frequency content and display settings.

内容的提问来源于stack exchange,提问作者Mattia Campana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 22:34:09