You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决NumPy数组处理WAV文件后播放噪音过大的问题?

问题

我有一段17秒的WAV文件,是Michael Bublé歌曲《Feeling Good》的结尾片段。我编写代码对该文件生成的NumPy数组进行傅里叶正变换与逆变换,绘制时域、频域图并播放音频;随后应用低通滤波器(阻断1000Hz以下频率),重复绘图与播放操作。但原音频及逆变换后的音频播放时噪音极大、音质极差,而滤波后的音频却几乎正常。我尝试了直接播放float32格式、按最大值归一化、转换为int16格式(乘以32767)、取绝对值等多种方式,均无法解决噪音问题,寻求解决方案。

完整代码

import numpy as np
import matplotlib.pyplot as plt
from scipy.io import wavfile
import sounddevice as sd    

# Load audio file
fs, x_t = wavfile.read('song.wav')
# If audio has only one channel, divide the signal by 2; otherwise, leave it as is
if x_t.ndim == 2:
    x_t = 0.5 * (x_t[:, 0] + x_t[:, 1])

n = len(x_t)
tiempo = n / fs  # Duration of the audio
Delta_t = 1 / fs  # Sampling time
x_t = x_t/np.max(np.abs(x_t))

# Plot time-domain representation of the audio
t = np.arange(0, n) * Delta_t
plt.subplot(3, 1, 1)
plt.plot(t, x_t)


# Compute and plot the spectrum
X_w = np.fft.fftshift(np.fft.fft(x_t))
Delta_f = 1 / (n * Delta_t)
f = np.arange(-n/2, n/2) * Delta_f
DE_X_w = np.abs(X_w)
plt.subplot(3, 1, 2)
plt.plot(f, DE_X_w / np.max(DE_X_w))

# Inverse transform and plot
x_t_inv = np.fft.ifft(np.fft.ifftshift(X_w))
x_t_2 = np.real(x_t_inv)
plt.subplot(3, 1, 3)
plt.plot(t, x_t_2)

print(max(x_t))
print(min(x_t))


print(max(x_t_2))
print(min(x_t_2))


scaled = np.int16(x_t_2 / np.max(np.abs(x_t_2)) * 32767)
#sd.play(scaled, fs)
"""
sd.play(x_t, fs)  # Convert to 16-bit integer format
sd.wait()
sd.play(x_t_2, fs)  # Convert to 16-bit integer format
sd.wait()
"""
sd.play(scaled, fs)  # Convert to 16-bit integer format
sd.wait()

# Low-pass filter
fpb = 1 * (np.abs(f) <= 1000)

# Plot the low-pass filter
plt.figure()
plt.subplot(3, 1, 1)
plt.plot(f, fpb)

# Apply the filter in the frequency domain
X_w_fil = X_w * fpb
plt.subplot(3, 1, 2)
plt.plot(f, np.abs(X_w_fil) / np.max(np.abs(X_w_fil)))

# Inverse transform and plot the filtered signal in the time domain
x_t_filt = np.fft.ifft(np.fft.ifftshift(X_w_fil))
x_t_filt = np.real(x_t_filt)
plt.subplot(3, 1, 3)
plt.plot(t, x_t_filt)

# Play the filtered audio
#sd.play((np.around(x_t_filt)).astype(np.int16), fs)  # Convert to 16-bit integer format
#sd.wait()


plt.show()
解决方案

1. 修复FFT逆变换的幅度归一化问题

NumPy的np.fft.fft和np.fft.ifft是未归一化的,逆变换结果的幅度是原信号的n倍(n为样本总数)。即使后续做了归一化,放大后的浮点误差会引入大量噪声,而滤波操作恰好滤除了这些高频噪声,所以滤波后音质正常。

修改逆变换代码,添加除以n的操作:

# 原逆变换代码
x_t_inv = np.fft.ifft(np.fft.ifftshift(X_w)) / n  # 新增除以n
x_t_2 = np.real(x_t_inv)

# 滤波后的逆变换代码
x_t_filt = np.fft.ifft(np.fft.ifftshift(X_w_fil)) / n  # 新增除以n
x_t_filt = np.real(x_t_filt)

2. 修复双声道处理的整数溢出问题

原代码中双声道平均时,若原音频是int16类型,两个int16值相加可能超出整数范围导致溢出,引入噪声。需先转换为浮点类型再计算:

if x_t.ndim == 2:
    # 先转浮点再相加,避免整数溢出
    x_t = 0.5 * (x_t[:, 0].astype(np.float64) + x_t[:, 1].astype(np.float64))

3. 规范音频播放的格式处理

sounddevice对浮点音频的要求是范围在[-1.0, 1.0],对int16音频要求范围在[-32767, 32767]。原信号已归一化到[-1,1],可直接播放float32格式,或正确转换为int16:

# 播放原音频(float32格式)
sd.play(x_t.astype(np.float32), fs)
sd.wait()

# 播放逆变换后的音频(正确转换int16)
scaled_inv = np.int16(x_t_2 / np.max(np.abs(x_t_2)) * 32767)
sd.play(scaled_inv, fs)
sd.wait()

4. 可选:减少浮点运算误差

逆变换后取实部时,可添加微小的阈值过滤浮点噪声:

x_t_2 = np.real(x_t_inv)
# 过滤小于1e-6的噪声
x_t_2[np.abs(x_t_2) < 1e-6] = 0

内容的提问来源于stack exchange,提问作者MrManini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 15:37:32