You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测WAV文件中蜂鸣声的频率与时长?修复短蜂鸣遗漏问题

修复WAV文件中短蜂鸣的频率与时长提取问题

我有一个test.wav文件,需要提取其中蜂鸣声的频率和时长。测试文件包含以下蜂鸣片段:

  • 1500 Hz 持续100 ms
  • 1900 Hz 持续300 ms
  • 1200 Hz 持续10 ms
  • 1900 Hz 持续300 ms

现有方案会遗漏所有短于100ms的蜂鸣,还存在频率检测偏高的问题,原代码如下:

import numpy as np
from scipy.io import wavfile

def extract_frequencies_and_durations(wav_path, frame_duration=0.02, freq_tolerance=5):
    rate, audio_data = wavfile.read(wav_path)
    if audio_data.ndim > 1:
        audio_data = np.mean(audio_data, axis=1)
    frame_length = int(rate * frame_duration)
    frequencies = []
    durations = []
    current_freq = None
    current_duration = 0
    for i in range(0, len(audio_data), frame_length):
        frame = audio_data[i:i + frame_length]
        fft_spectrum = np.fft.fft(frame)
        freqs = np.fft.fftfreq(len(fft_spectrum), 1 / rate)
        magnitude = np.abs(fft_spectrum)
        dominant_freq = abs(freqs[np.argmax(magnitude)])
        if current_freq is None or abs(dominant_freq - current_freq) > freq_tolerance:
            if current_freq is not None:
                frequencies.append(current_freq)
                durations.append(current_duration)
            current_freq = dominant_freq
            current_duration = frame_duration
        else:
            current_duration += frame_duration
    if current_freq is not None:
        frequencies.append(current_freq)
        durations.append(current_duration)
    
    return frequencies, durations

# output
wav_path = 'test.wav'
frequencies, durations = extract_frequencies_and_durations(wav_path)
for freq, dur in zip(frequencies, durations):
    print(f"{freq:.2f}Hz - {dur * 1000:.2f}ms")

原代码核心问题

  1. 非重叠固定帧长:20ms帧长且步进等于帧长,10ms的短蜂鸣仅覆盖部分帧,易被前后帧信号淹没,无法被识别为独立段。
  2. 频谱泄漏与分辨率不足:未加窗导致FFT频谱泄漏,直接取峰值bin的频率会产生偏移;固定帧长的频率分辨率(采样率/帧长)偏低,无法精准定位真实频率。
  3. 简单频率判定逻辑:仅通过单帧频率与当前频率的差值判断切换,短帧信号易受相邻帧干扰,导致漏检。

修复后的代码

import numpy as np
from scipy.io import wavfile

def extract_frequencies_and_durations(wav_path, frame_duration=0.005, hop_duration=0.001, freq_tolerance=10, min_segment_duration=0.005):
    # 读取音频并转单声道
    rate, audio_data = wavfile.read(wav_path)
    if audio_data.ndim > 1:
        audio_data = np.mean(audio_data, axis=1).astype(np.int16)
    # 归一化处理
    audio_data = audio_data / np.iinfo(np.int16).max
    
    frame_length = int(rate * frame_duration)
    hop_length = int(rate * hop_duration)
    # 预加汉宁窗减少频谱泄漏
    window = np.hanning(frame_length)
    
    frequencies = []
    durations = []
    current_freq = None
    current_duration = 0
    current_segments = []
    
    # 遍历重叠帧
    for i in range(0, len(audio_data) - frame_length + 1, hop_length):
        frame = audio_data[i:i+frame_length] * window
        
        # 计算FFT并提取正频率部分
        fft_spectrum = np.fft.fft(frame)[:frame_length//2]
        freqs = np.fft.fftfreq(frame_length, 1/rate)[:frame_length//2]
        magnitude = np.abs(fft_spectrum)
        
        # 抛物线插值提升频率精度
        peak_idx = np.argmax(magnitude)
        if 0 < peak_idx < len(magnitude)-1:
            left, mid, right = magnitude[peak_idx-1], magnitude[peak_idx], magnitude[peak_idx+1]
            delta = (right - left) / (2 * (2*mid - left - right))
            dominant_freq = freqs[peak_idx] + delta * (freqs[1] - freqs[0])
        else:
            dominant_freq = freqs[peak_idx]
        
        # 基于平均频率的切换判定
        if current_freq is None:
            current_freq = dominant_freq
            current_duration = hop_duration
            current_segments.append(dominant_freq)
        else:
            avg_current = np.mean(current_segments)
            if abs(dominant_freq - avg_current) > freq_tolerance:
                # 保存有效段(过滤过短噪声)
                if current_duration >= min_segment_duration:
                    frequencies.append(avg_current)
                    durations.append(current_duration)
                # 切换到新频率段
                current_freq = dominant_freq
                current_duration = hop_duration
                current_segments = [dominant_freq]
            else:
                current_duration += hop_duration
                current_segments.append(dominant_freq)
    
    # 处理最后一段
    if current_freq is not None and current_duration >= min_segment_duration:
        frequencies.append(np.mean(current_segments))
        durations.append(current_duration)
    
    return frequencies, durations

# 测试执行
wav_path = 'test.wav'
frequencies, durations = extract_frequencies_and_durations(wav_path)
for freq, dur in zip(frequencies, durations):
    print(f"{round(freq)}Hz - {dur * 1000:.1f}ms")

关键改进点

  • 重叠短帧:使用5ms帧长+1ms步进的重叠帧,确保10ms短蜂鸣能被多帧捕获,避免漏检。
  • 频谱优化:汉宁窗减少泄漏,抛物线插值提升频率检测精度,解决频率偏高问题。
  • 稳定判定逻辑:基于当前段的平均频率判断切换,避免单帧噪声干扰,提升短信号识别稳定性。
  • 灵活阈值:可自定义最小检测时长,过滤无效噪声同时保留有效短信号。

内容的提问来源于stack exchange,提问作者kotek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 13:52:03