如何解决NumPy数组处理WAV文件后播放噪音过大的问题?
问题
我有一段17秒的WAV文件,是Michael Bublé歌曲《Feeling Good》的结尾片段。我编写代码对该文件生成的NumPy数组进行傅里叶正变换与逆变换,绘制时域、频域图并播放音频;随后应用低通滤波器(阻断1000Hz以下频率),重复绘图与播放操作。但原音频及逆变换后的音频播放时噪音极大、音质极差,而滤波后的音频却几乎正常。我尝试了直接播放float32格式、按最大值归一化、转换为int16格式(乘以32767)、取绝对值等多种方式,均无法解决噪音问题,寻求解决方案。
完整代码
import numpy as np import matplotlib.pyplot as plt from scipy.io import wavfile import sounddevice as sd # Load audio file fs, x_t = wavfile.read('song.wav') # If audio has only one channel, divide the signal by 2; otherwise, leave it as is if x_t.ndim == 2: x_t = 0.5 * (x_t[:, 0] + x_t[:, 1]) n = len(x_t) tiempo = n / fs # Duration of the audio Delta_t = 1 / fs # Sampling time x_t = x_t/np.max(np.abs(x_t)) # Plot time-domain representation of the audio t = np.arange(0, n) * Delta_t plt.subplot(3, 1, 1) plt.plot(t, x_t) # Compute and plot the spectrum X_w = np.fft.fftshift(np.fft.fft(x_t)) Delta_f = 1 / (n * Delta_t) f = np.arange(-n/2, n/2) * Delta_f DE_X_w = np.abs(X_w) plt.subplot(3, 1, 2) plt.plot(f, DE_X_w / np.max(DE_X_w)) # Inverse transform and plot x_t_inv = np.fft.ifft(np.fft.ifftshift(X_w)) x_t_2 = np.real(x_t_inv) plt.subplot(3, 1, 3) plt.plot(t, x_t_2) print(max(x_t)) print(min(x_t)) print(max(x_t_2)) print(min(x_t_2)) scaled = np.int16(x_t_2 / np.max(np.abs(x_t_2)) * 32767) #sd.play(scaled, fs) """ sd.play(x_t, fs) # Convert to 16-bit integer format sd.wait() sd.play(x_t_2, fs) # Convert to 16-bit integer format sd.wait() """ sd.play(scaled, fs) # Convert to 16-bit integer format sd.wait() # Low-pass filter fpb = 1 * (np.abs(f) <= 1000) # Plot the low-pass filter plt.figure() plt.subplot(3, 1, 1) plt.plot(f, fpb) # Apply the filter in the frequency domain X_w_fil = X_w * fpb plt.subplot(3, 1, 2) plt.plot(f, np.abs(X_w_fil) / np.max(np.abs(X_w_fil))) # Inverse transform and plot the filtered signal in the time domain x_t_filt = np.fft.ifft(np.fft.ifftshift(X_w_fil)) x_t_filt = np.real(x_t_filt) plt.subplot(3, 1, 3) plt.plot(t, x_t_filt) # Play the filtered audio #sd.play((np.around(x_t_filt)).astype(np.int16), fs) # Convert to 16-bit integer format #sd.wait() plt.show()
解决方案
1. 修复FFT逆变换的幅度归一化问题
NumPy的np.fft.fft和np.fft.ifft是未归一化的,逆变换结果的幅度是原信号的n倍(n为样本总数)。即使后续做了归一化,放大后的浮点误差会引入大量噪声,而滤波操作恰好滤除了这些高频噪声,所以滤波后音质正常。
修改逆变换代码,添加除以n的操作:
# 原逆变换代码 x_t_inv = np.fft.ifft(np.fft.ifftshift(X_w)) / n # 新增除以n x_t_2 = np.real(x_t_inv) # 滤波后的逆变换代码 x_t_filt = np.fft.ifft(np.fft.ifftshift(X_w_fil)) / n # 新增除以n x_t_filt = np.real(x_t_filt)
2. 修复双声道处理的整数溢出问题
原代码中双声道平均时,若原音频是int16类型,两个int16值相加可能超出整数范围导致溢出,引入噪声。需先转换为浮点类型再计算:
if x_t.ndim == 2: # 先转浮点再相加,避免整数溢出 x_t = 0.5 * (x_t[:, 0].astype(np.float64) + x_t[:, 1].astype(np.float64))
3. 规范音频播放的格式处理
sounddevice对浮点音频的要求是范围在[-1.0, 1.0],对int16音频要求范围在[-32767, 32767]。原信号已归一化到[-1,1],可直接播放float32格式,或正确转换为int16:
# 播放原音频(float32格式) sd.play(x_t.astype(np.float32), fs) sd.wait() # 播放逆变换后的音频(正确转换int16) scaled_inv = np.int16(x_t_2 / np.max(np.abs(x_t_2)) * 32767) sd.play(scaled_inv, fs) sd.wait()
4. 可选:减少浮点运算误差
逆变换后取实部时,可添加微小的阈值过滤浮点噪声:
x_t_2 = np.real(x_t_inv) # 过滤小于1e-6的噪声 x_t_2[np.abs(x_t_2) < 1e-6] = 0
内容的提问来源于stack exchange,提问作者MrManini
相关产品推荐
相关产品推荐

