You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

wave.readframes()与librosa.load()读取音频数据的差异原因咨询

Why do librosa.load() and wave.readframes() return drastically different values for the same WAV file?

Great question! The difference you're seeing comes down to how each library processes and represents audio data—let's break it down clearly:

1. Automatic Normalization (The Core Reason)

Librosa's load() function automatically normalizes audio data to the [-1.0, 1.0] range by default. This is standard for audio processing libraries because floating-point values in this range are far easier to work with for operations like STFT, feature extraction, or neural network inputs.

Your WAV file uses 16-bit PCM encoding (the most common format), which stores raw sample values as integers between -32768 and 32767. Librosa takes these raw integers and divides each by 32768 to scale them into the [-1, 1] float range.

Looking at your output, you can see this exact relationship:

  • Wave's first value is -1 → -1 / 32768 = -3.0517578e-05 (matches Librosa's first value)
  • Wave's second value is 13 → 13 / 32768 ≈ 3.9672852e-04 (matches Librosa's second value)

2. Data Type Differences

  • wave.readframes() returns raw byte data, which you converted to np.int16—this is the original integer format the WAV file stores samples in.
  • librosa.load() returns float32 values by default, even if you disable normalization (using normalize=False). Floating-point arithmetic is standard for audio processing due to better precision during complex operations.

3. Extra Librosa Processing (No Impact on Your Sample Length)

Librosa handles a few more things behind the scenes that didn't affect your sample count here, but are useful to know:

  • It converts stereo audio to mono by default (use mono=False to preserve channels). Your length matches wave's output, so your file is already mono.
  • It resamples the audio to your specified sr=44100 if the original sample rate differs. Since your length matches, your file's native sample rate is already 44100.

Quick Verification Code

To confirm this relationship, run this snippet—it scales Librosa's data back to the original integer range and checks if it matches the wave output:

import numpy as np
import librosa
import wave

sample_wave = './data/mywave.wav'

# Load with Librosa
a, sr = librosa.load(sample_wave, sr=44100)
# Load with wave
wav = wave.open(sample_wave)
data = wav.readframes(wav.getnframes())
b = np.frombuffer(data, dtype=np.int16)

# Scale Librosa's float data back to int16
a_scaled = (a * 32768).astype(np.int16)

# Check if they match
print(np.array_equal(a_scaled, b))  # Should print True!

内容的提问来源于stack exchange,提问作者whitebear

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 20:18:10