Python读取WAV文件:采样率与数据维度不匹配及波形频谱图生成需求
Hey there! Let's work through that dimension mismatch issue you're facing when generating the spectrogram. Here's a breakdown of what's likely going on and how to fix it:
First, Let's Diagnose the Problem
From your output, I can see your samples variable is a 1-dimensional array (shape: (14466512,)), which makes sense for a mono WAV file. Most spectrogram functions (custom or library-provided) expect either:
- A 2-dimensional array (shape:
(number_of_samples, number_of_channels)) for multi-channel audio, or - A 1-dimensional array specifically for mono audio, but sometimes functions are built to handle multi-channel inputs by default.
The "dimension mismatch" error is almost certainly because your log_specgram function is expecting a 2D array, but you're passing a 1D array.
Step-by-Step Fixes
1. Check and Adjust the Sample Array Dimension
First, confirm the shape of your samples to be sure:
print(samples.shape) # Should output (14466512,) for mono audio
If it's 1D, convert it to a 2D array (since mono audio is just 1 channel):
import numpy as np # Reshape to (number_of_samples, 1) for single-channel audio samples = samples.reshape(-1, 1)
2. Convert to Float and Normalize (Critical for Most Spectrogram Functions)
wavfile.read() returns integer values (usually int16), but many spectral analysis functions expect floating-point data normalized to the range [-1, 1]. Fix this with:
# Convert to float32 and normalize samples = samples.astype(np.float32) / np.iinfo(np.int16).max
3. Verify Function Parameter Order
Double-check that your log_specgram function expects (samples, sample_rate) as inputs. Some functions reverse the order (e.g., (sample_rate, samples)). If you're using a custom log_specgram, confirm its definition matches how you're calling it.
Full Working Example (Waveform + Spectrogram)
Here's a complete code snippet that includes both plots, assuming you're using standard libraries like numpy, scipy, and matplotlib:
import numpy as np from scipy.io import wavfile from scipy import signal import matplotlib.pyplot as plt # Custom log_specgram implementation (adjust if yours is different) def log_specgram(audio, sample_rate, window_size=20, step_size=10, eps=1e-10): nperseg = int(round(window_size * sample_rate / 1000)) noverlap = int(round(step_size * sample_rate / 1000)) freqs, times, spec = signal.spectrogram( audio, fs=sample_rate, nperseg=nperseg, noverlap=noverlap ) return freqs, times, np.log(spec.T.astype(np.float32) + eps) # Load your audio file sample_rate, samples = wavfile.read(files[:1][0]) print(f"Sample Rate: {sample_rate}") print(f"Original Samples Shape: {samples.shape}") # Fix dimension and data type if len(samples.shape) == 1: samples = samples.reshape(-1, 1) samples = samples.astype(np.float32) / np.iinfo(np.int16).max # Generate spectrogram using the first (and only) channel freqs, times, spectrogram = log_specgram(samples[:, 0], sample_rate) # Plot Waveform plt.figure(figsize=(12, 8)) plt.subplot(2, 1, 1) plt.plot(np.arange(len(samples)) / sample_rate, samples[:, 0]) plt.title("Audio Waveform") plt.xlabel("Time (seconds)") plt.ylabel("Amplitude") # Plot Spectrogram plt.subplot(2, 1, 2) plt.imshow( spectrogram.T, aspect="auto", origin="lower", extent=[times.min(), times.max(), freqs.min(), freqs.max()] ) plt.title("Log Spectrogram") plt.xlabel("Time (seconds)") plt.ylabel("Frequency (Hz)") plt.colorbar(label="Log Amplitude") plt.tight_layout() plt.show()
Additional Notes
- If your
log_specgramis from a library likelibrosa, it likely expects a 1D array for mono audio. In that case, skip thereshapestep and just passsamplesdirectly (after normalizing to float). - If you still get errors, print the full traceback—it will tell you exactly which dimension the function is expecting vs. what you're passing.
内容的提问来源于stack exchange,提问作者bytebiscuit

