You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

傅里叶变换基音检测函数中3个局部最大值识别错误,如何修正?

钢琴基频识别错误:E4误判为B5的问题排查与修复

问题描述

我实现了一个麦克风输入流的频率检测函数,用于提取三个最显著频率,但存在严重的基频误判问题:

  • 弹奏钢琴E4(392 Hz)时,函数返回的基频是B5(996 Hz)
  • 偶尔会将C4识别为C5,其中E4的错误表现最突出
  • 观察频谱图谱(下图),E4的基频峰值明明是最显著的,但函数仍会误选谐波作为基频

弹奏E4音符的傅里叶图谱

原始代码

def pitch_calculations(stream, CHUNK, RATE):
    # Read mic stream and then call struct.unpack to convert from binary data back to floats
    data = stream.read(CHUNK, exception_on_overflow=False)
    dataInt = np.array(struct.unpack(str(CHUNK) + 'h', data))
    
    # Apply a window function (Hamming) to the input data
    windowed_data = np.hamming(CHUNK) * dataInt

    # Using numpy fast Fourier transform to convert mic data into frequencies
    fft_result = np.abs(np.fft.fft(windowed_data)) * 2 / (11000 * CHUNK)
    freqs = np.fft.fftfreq(len(windowed_data), d=1.0 / RATE)
    
    # Find the indices of local maxima in the frequency spectrum
    localmax_indecies = argrelextrema(fft_result, np.greater)[0]
    
    # Get the magnitudes of the local maxima
    strong_freqs = fft_result[localmax_indecies]
    
    # Sort the magnitudes in descending order
    sorted_indices = np.argsort(strong_freqs)[::-1]
    
    # Get the indices of the three highest peaks
    top_indices = sorted_indices[:6]
    
    # Get the corresponding frequencies
    note_1_freq = abs(freqs[localmax_indecies[top_indices[0]]])
    note_2_freq = abs(freqs[localmax_indecies[top_indices[2]]])
    note_3_freq = abs(freqs[localmax_indecies[top_indices[4]]])
    
    return note_1_freq, note_2_freq, note_3_freq

问题根源分析

  1. 错误的峰值选取逻辑:代码中取前6个排序后的峰值索引,然后跳着选top_indices[0]、top_indices[2]、top_indices[4],完全违背了“取最显著三个频率”的需求,这种无意义的跳跃直接导致高幅值谐波被优先选中。
  2. FFT幅度归一化错误:硬编码的11000并非标准16位PCM的最大幅值(应为32767),导致幅度计算失真,可能影响峰值排序的准确性。
  3. 未区分基频与谐波:钢琴音符的谐波(整数倍频率)幅值常高于基频,但基频是最低的有效频率。当前代码仅按幅值排序,必然会误把谐波当作基频。
  4. 频率分辨率不足:若RATE/CHUNK的比值过大,频谱分辨率会过低,无法精准区分基频与相邻谐波。

修复后的代码

import numpy as np
import struct
from scipy.signal import argrelextrema

def pitch_calculations(stream, CHUNK, RATE):
    # 读取麦克风数据并转换为整数数组
    data = stream.read(CHUNK, exception_on_overflow=False)
    dataInt = np.array(struct.unpack(str(CHUNK) + 'h', data))
    
    # 应用汉明窗减少频谱泄漏
    windowed_data = np.hamming(CHUNK) * dataInt

    # FFT计算并归一化幅度(16位PCM最大幅值为32767)
    fft_result = np.abs(np.fft.fft(windowed_data)) * 2 / (32767 * CHUNK)
    freqs = np.fft.fftfreq(len(windowed_data), d=1.0 / RATE)
    # 只保留正频率部分,简化处理
    positive_mask = freqs > 0
    freqs = freqs[positive_mask]
    fft_result = fft_result[positive_mask]
    
    # 查找局部峰值,并过滤低幅值噪声
    localmax_indices = argrelextrema(fft_result, np.greater)[0]
    min_magnitude = 0.01  # 根据实际环境调整阈值
    valid_peaks = [(freqs[i], fft_result[i]) for i in localmax_indices if fft_result[i] > min_magnitude]
    if not valid_peaks:
        return None, None, None
    
    # 按幅值降序排序候选峰值
    valid_peaks.sort(key=lambda x: x[1], reverse=True)
    
    # 基频检测逻辑:寻找能匹配多个谐波的最低频率
    def is_fundamental(candidate, others, tolerance=0.03):
        # 检查候选频率是否存在多个整数倍的谐波峰值
        for freq in others:
            ratio = freq / candidate
            if abs(ratio - round(ratio)) < tolerance:
                return True
        return False
    
    # 筛选基频
    fundamental = None
    top_candidates = [peak[0] for peak in valid_peaks[:10]]
    for i, candidate in enumerate(top_candidates):
        others = top_candidates[:i] + top_candidates[i+1:]
        if is_fundamental(candidate, others):
            fundamental = candidate
            break
    
    # 构建结果列表:优先基频,再取非谐波的显著频率
    result = []
    if fundamental:
        result.append(fundamental)
        # 过滤基频的谐波,取剩余幅值最高的两个
        harmonic_tolerance = 0.03
        non_harmonic_peaks = []
        for peak in valid_peaks[1:]:
            ratio = peak[0] / fundamental
            if abs(ratio - round(ratio)) > harmonic_tolerance:
                non_harmonic_peaks.append(peak[0])
        # 非谐波不足时补充谐波峰值
        result.extend(non_harmonic_peaks[:2])
        if len(result) < 3:
            result.extend([peak[0] for peak in valid_peaks[1:(3 - len(result) + 1)]])
    else:
        # 无法检测基频时,直接取前三个幅值最高的频率
        result = [peak[0] for peak in valid_peaks[:3]]
    
    # 补全三个结果位
    while len(result) < 3:
        result.append(None)
    
    return result[0], result[1], result[2]

额外优化建议

  • 提升频率分辨率:增大CHUNK值(比如从1024改为4096),降低RATE/CHUNK的比值,让频谱更精细,减少基频与谐波的重叠误差。
  • 峰值精细化:使用抛物线插值对峰值频率进行修正,进一步提高频率检测精度。
  • 多算法融合:结合自相关法(如YIN算法)验证FFT得到的基频,避免单一算法的局限性。

内容的提问来源于stack exchange,提问作者celewis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 06:22:35