傅里叶变换基音检测函数中3个局部最大值识别错误,如何修正?
钢琴基频识别错误:E4误判为B5的问题排查与修复
问题描述
我实现了一个麦克风输入流的频率检测函数,用于提取三个最显著频率,但存在严重的基频误判问题:
- 弹奏钢琴E4(392 Hz)时,函数返回的基频是B5(996 Hz)
- 偶尔会将C4识别为C5,其中E4的错误表现最突出
- 观察频谱图谱(下图),E4的基频峰值明明是最显著的,但函数仍会误选谐波作为基频

原始代码
def pitch_calculations(stream, CHUNK, RATE): # Read mic stream and then call struct.unpack to convert from binary data back to floats data = stream.read(CHUNK, exception_on_overflow=False) dataInt = np.array(struct.unpack(str(CHUNK) + 'h', data)) # Apply a window function (Hamming) to the input data windowed_data = np.hamming(CHUNK) * dataInt # Using numpy fast Fourier transform to convert mic data into frequencies fft_result = np.abs(np.fft.fft(windowed_data)) * 2 / (11000 * CHUNK) freqs = np.fft.fftfreq(len(windowed_data), d=1.0 / RATE) # Find the indices of local maxima in the frequency spectrum localmax_indecies = argrelextrema(fft_result, np.greater)[0] # Get the magnitudes of the local maxima strong_freqs = fft_result[localmax_indecies] # Sort the magnitudes in descending order sorted_indices = np.argsort(strong_freqs)[::-1] # Get the indices of the three highest peaks top_indices = sorted_indices[:6] # Get the corresponding frequencies note_1_freq = abs(freqs[localmax_indecies[top_indices[0]]]) note_2_freq = abs(freqs[localmax_indecies[top_indices[2]]]) note_3_freq = abs(freqs[localmax_indecies[top_indices[4]]]) return note_1_freq, note_2_freq, note_3_freq
问题根源分析
- 错误的峰值选取逻辑:代码中取前6个排序后的峰值索引,然后跳着选
top_indices[0]、top_indices[2]、top_indices[4],完全违背了“取最显著三个频率”的需求,这种无意义的跳跃直接导致高幅值谐波被优先选中。 - FFT幅度归一化错误:硬编码的
11000并非标准16位PCM的最大幅值(应为32767),导致幅度计算失真,可能影响峰值排序的准确性。 - 未区分基频与谐波:钢琴音符的谐波(整数倍频率)幅值常高于基频,但基频是最低的有效频率。当前代码仅按幅值排序,必然会误把谐波当作基频。
- 频率分辨率不足:若
RATE/CHUNK的比值过大,频谱分辨率会过低,无法精准区分基频与相邻谐波。
修复后的代码
import numpy as np import struct from scipy.signal import argrelextrema def pitch_calculations(stream, CHUNK, RATE): # 读取麦克风数据并转换为整数数组 data = stream.read(CHUNK, exception_on_overflow=False) dataInt = np.array(struct.unpack(str(CHUNK) + 'h', data)) # 应用汉明窗减少频谱泄漏 windowed_data = np.hamming(CHUNK) * dataInt # FFT计算并归一化幅度(16位PCM最大幅值为32767) fft_result = np.abs(np.fft.fft(windowed_data)) * 2 / (32767 * CHUNK) freqs = np.fft.fftfreq(len(windowed_data), d=1.0 / RATE) # 只保留正频率部分,简化处理 positive_mask = freqs > 0 freqs = freqs[positive_mask] fft_result = fft_result[positive_mask] # 查找局部峰值,并过滤低幅值噪声 localmax_indices = argrelextrema(fft_result, np.greater)[0] min_magnitude = 0.01 # 根据实际环境调整阈值 valid_peaks = [(freqs[i], fft_result[i]) for i in localmax_indices if fft_result[i] > min_magnitude] if not valid_peaks: return None, None, None # 按幅值降序排序候选峰值 valid_peaks.sort(key=lambda x: x[1], reverse=True) # 基频检测逻辑:寻找能匹配多个谐波的最低频率 def is_fundamental(candidate, others, tolerance=0.03): # 检查候选频率是否存在多个整数倍的谐波峰值 for freq in others: ratio = freq / candidate if abs(ratio - round(ratio)) < tolerance: return True return False # 筛选基频 fundamental = None top_candidates = [peak[0] for peak in valid_peaks[:10]] for i, candidate in enumerate(top_candidates): others = top_candidates[:i] + top_candidates[i+1:] if is_fundamental(candidate, others): fundamental = candidate break # 构建结果列表:优先基频,再取非谐波的显著频率 result = [] if fundamental: result.append(fundamental) # 过滤基频的谐波,取剩余幅值最高的两个 harmonic_tolerance = 0.03 non_harmonic_peaks = [] for peak in valid_peaks[1:]: ratio = peak[0] / fundamental if abs(ratio - round(ratio)) > harmonic_tolerance: non_harmonic_peaks.append(peak[0]) # 非谐波不足时补充谐波峰值 result.extend(non_harmonic_peaks[:2]) if len(result) < 3: result.extend([peak[0] for peak in valid_peaks[1:(3 - len(result) + 1)]]) else: # 无法检测基频时,直接取前三个幅值最高的频率 result = [peak[0] for peak in valid_peaks[:3]] # 补全三个结果位 while len(result) < 3: result.append(None) return result[0], result[1], result[2]
额外优化建议
- 提升频率分辨率:增大
CHUNK值(比如从1024改为4096),降低RATE/CHUNK的比值,让频谱更精细,减少基频与谐波的重叠误差。 - 峰值精细化:使用抛物线插值对峰值频率进行修正,进一步提高频率检测精度。
- 多算法融合:结合自相关法(如YIN算法)验证FFT得到的基频,避免单一算法的局限性。
内容的提问来源于stack exchange,提问作者celewis
相关产品推荐
相关产品推荐

