wave.readframes()与librosa.load()读取音频数据的差异原因咨询
Great question! The difference you're seeing comes down to how each library processes and represents audio data—let's break it down clearly:
1. Automatic Normalization (The Core Reason)
Librosa's load() function automatically normalizes audio data to the [-1.0, 1.0] range by default. This is standard for audio processing libraries because floating-point values in this range are far easier to work with for operations like STFT, feature extraction, or neural network inputs.
Your WAV file uses 16-bit PCM encoding (the most common format), which stores raw sample values as integers between -32768 and 32767. Librosa takes these raw integers and divides each by 32768 to scale them into the [-1, 1] float range.
Looking at your output, you can see this exact relationship:
- Wave's first value is
-1→-1 / 32768 = -3.0517578e-05(matches Librosa's first value) - Wave's second value is
13→13 / 32768 ≈ 3.9672852e-04(matches Librosa's second value)
2. Data Type Differences
wave.readframes()returns raw byte data, which you converted tonp.int16—this is the original integer format the WAV file stores samples in.librosa.load()returnsfloat32values by default, even if you disable normalization (usingnormalize=False). Floating-point arithmetic is standard for audio processing due to better precision during complex operations.
3. Extra Librosa Processing (No Impact on Your Sample Length)
Librosa handles a few more things behind the scenes that didn't affect your sample count here, but are useful to know:
- It converts stereo audio to mono by default (use
mono=Falseto preserve channels). Your length matches wave's output, so your file is already mono. - It resamples the audio to your specified
sr=44100if the original sample rate differs. Since your length matches, your file's native sample rate is already 44100.
Quick Verification Code
To confirm this relationship, run this snippet—it scales Librosa's data back to the original integer range and checks if it matches the wave output:
import numpy as np import librosa import wave sample_wave = './data/mywave.wav' # Load with Librosa a, sr = librosa.load(sample_wave, sr=44100) # Load with wave wav = wave.open(sample_wave) data = wav.readframes(wav.getnframes()) b = np.frombuffer(data, dtype=np.int16) # Scale Librosa's float data back to int16 a_scaled = (a * 32768).astype(np.int16) # Check if they match print(np.array_equal(a_scaled, b)) # Should print True!
内容的提问来源于stack exchange,提问作者whitebear

