You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python计算.wav文件频谱图异常问题排查求助

Troubleshooting Your Wav Spectrogram Issues in Python

Hey there, let's work through both of your spectrogram problems one by one—first the weird blue stripes in your custom sig2spec function, then the all-dark-blue output from librosa's built-in melspectrogram.

Problem 1: Blue Stripes at the Start of Custom sig2spec Output

Your hunch about normalization is spot-on—let's break down why those stripes are happening:

Root Cause

Looking at your final line of sig2spec:

return (filter_banks/ np.amax(filter_banks))*255

After converting to dB with 20 * numpy.log10(filter_banks), your filter_banks values are mostly negative (since dB scales log power, and most frames will have lower power than the loudest frame). When you divide by the global maximum (which could be 0 dB for the loudest frame), negative values get scaled to negative numbers, which get clamped to 0 when rendering—resulting in deep blue (the lowest end of the jet colormap).

Additionally, the pre-emphasis step might amplify tiny artifacts in the first frame: you're keeping signal[0] as-is, then applying signal[1:] - pre_emphasis * signal[:-1] to the rest. This could create a sharp contrast between the first frame and subsequent ones.

Fixes

  1. Fix Normalization for dB Values: Instead of dividing by the global max (which doesn't handle negative dB properly), normalize the entire range to 0-255:
    # Replace your return line with this
    filter_banks -= np.min(filter_banks)  # Shift minimum value to 0
    filter_banks /= np.max(filter_banks)  # Scale to 0-1 range
    return filter_banks * 255
    
  2. Test Pre-Emphasis Impact: Temporarily comment out the pre-emphasis step to see if the stripes disappear. If they do, adjust the pre_emphasis value (try 0.95 instead of 0.97) or add a small fade-in to the start of your signal to smooth the first frame.

Problem 2: All-Dark Blue Output from Librosa's melspectrogram

Your librosa code is missing a critical step—converting linear power to dB scale—and has a couple of minor formatting issues:

Root Causes

  1. No dB Conversion: librosa.feature.melspectrogram returns linear power values, which have an enormous dynamic range. Most of your signal's power will be clustered at the low end of this range, rendering as deep blue in the jet colormap.
  2. Incorrect Spectrogram Transposition: Librosa's melspectrogram outputs shape (n_mels, time_frames), and librosa.display.specshow expects this format—your .T transpose flips the axes, which distorts the visualization.
  3. Unrealistic Figure Size: Setting figsize=(spec_shape) creates an absurdly large figure (e.g., if your spectrogram is (40, 1000), the figure would be 40x1000 inches!), leading to a squashed, unreadable output.

Fixed Librosa Code

import librosa
import librosa.display
import matplotlib.pyplot as plt
import numpy as np

# Load audio without resampling
sig, rate = librosa.load("audio.wav", sr=None)

# Generate mel spectrogram (linear power)
spectrogram = librosa.feature.melspectrogram(y=sig, sr=rate, n_mels=40)  # Match nfilt=40 from your custom function

# Convert to dB scale (critical for human-readable visualization)
spectrogram_db = librosa.power_to_db(spectrogram, ref=np.max)  # Normalize to loudest frame as 0 dB

# Create a properly sized figure
fig, ax = plt.subplots(figsize=(12, 4), dpi=100)

# Display spectrogram with correct axes
img = librosa.display.specshow(
    spectrogram_db,
    sr=rate,
    x_axis='time',
    y_axis='mel',
    cmap='jet',
    ax=ax
)

# Add colorbar for dB reference
fig.colorbar(img, ax=ax, format='%+2.0f dB')
plt.title('Mel Spectrogram (dB Scale)')
plt.tight_layout()
plt.savefig("spec.jpg")
plt.close()

Key Changes Explained

  • librosa.power_to_db: Converts linear power to logarithmic dB scale, which aligns with how humans perceive sound and compresses the dynamic range so you can see detail across quiet and loud sections.
  • Removed .T: Keeps the spectrogram in the shape librosa.display.specshow expects.
  • Reasonable figsize: Ensures the output image is readable and properly scaled.

内容的提问来源于stack exchange,提问作者Jose Ramon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:20:45