Python计算.wav文件频谱图异常问题排查求助
Hey there, let's work through both of your spectrogram problems one by one—first the weird blue stripes in your custom sig2spec function, then the all-dark-blue output from librosa's built-in melspectrogram.
Problem 1: Blue Stripes at the Start of Custom sig2spec Output
Your hunch about normalization is spot-on—let's break down why those stripes are happening:
Root Cause
Looking at your final line of sig2spec:
return (filter_banks/ np.amax(filter_banks))*255
After converting to dB with 20 * numpy.log10(filter_banks), your filter_banks values are mostly negative (since dB scales log power, and most frames will have lower power than the loudest frame). When you divide by the global maximum (which could be 0 dB for the loudest frame), negative values get scaled to negative numbers, which get clamped to 0 when rendering—resulting in deep blue (the lowest end of the jet colormap).
Additionally, the pre-emphasis step might amplify tiny artifacts in the first frame: you're keeping signal[0] as-is, then applying signal[1:] - pre_emphasis * signal[:-1] to the rest. This could create a sharp contrast between the first frame and subsequent ones.
Fixes
- Fix Normalization for dB Values: Instead of dividing by the global max (which doesn't handle negative dB properly), normalize the entire range to 0-255:
# Replace your return line with this filter_banks -= np.min(filter_banks) # Shift minimum value to 0 filter_banks /= np.max(filter_banks) # Scale to 0-1 range return filter_banks * 255 - Test Pre-Emphasis Impact: Temporarily comment out the pre-emphasis step to see if the stripes disappear. If they do, adjust the
pre_emphasisvalue (try 0.95 instead of 0.97) or add a small fade-in to the start of your signal to smooth the first frame.
Problem 2: All-Dark Blue Output from Librosa's melspectrogram
Your librosa code is missing a critical step—converting linear power to dB scale—and has a couple of minor formatting issues:
Root Causes
- No dB Conversion:
librosa.feature.melspectrogramreturns linear power values, which have an enormous dynamic range. Most of your signal's power will be clustered at the low end of this range, rendering as deep blue in thejetcolormap. - Incorrect Spectrogram Transposition: Librosa's
melspectrogramoutputs shape(n_mels, time_frames), andlibrosa.display.specshowexpects this format—your.Ttranspose flips the axes, which distorts the visualization. - Unrealistic Figure Size: Setting
figsize=(spec_shape)creates an absurdly large figure (e.g., if your spectrogram is (40, 1000), the figure would be 40x1000 inches!), leading to a squashed, unreadable output.
Fixed Librosa Code
import librosa import librosa.display import matplotlib.pyplot as plt import numpy as np # Load audio without resampling sig, rate = librosa.load("audio.wav", sr=None) # Generate mel spectrogram (linear power) spectrogram = librosa.feature.melspectrogram(y=sig, sr=rate, n_mels=40) # Match nfilt=40 from your custom function # Convert to dB scale (critical for human-readable visualization) spectrogram_db = librosa.power_to_db(spectrogram, ref=np.max) # Normalize to loudest frame as 0 dB # Create a properly sized figure fig, ax = plt.subplots(figsize=(12, 4), dpi=100) # Display spectrogram with correct axes img = librosa.display.specshow( spectrogram_db, sr=rate, x_axis='time', y_axis='mel', cmap='jet', ax=ax ) # Add colorbar for dB reference fig.colorbar(img, ax=ax, format='%+2.0f dB') plt.title('Mel Spectrogram (dB Scale)') plt.tight_layout() plt.savefig("spec.jpg") plt.close()
Key Changes Explained
librosa.power_to_db: Converts linear power to logarithmic dB scale, which aligns with how humans perceive sound and compresses the dynamic range so you can see detail across quiet and loud sections.- Removed
.T: Keeps the spectrogram in the shapelibrosa.display.specshowexpects. - Reasonable
figsize: Ensures the output image is readable and properly scaled.
内容的提问来源于stack exchange,提问作者Jose Ramon

