如何检测WAV文件中蜂鸣声的频率与时长?修复短蜂鸣遗漏问题
修复WAV文件中短蜂鸣的频率与时长提取问题
我有一个test.wav文件,需要提取其中蜂鸣声的频率和时长。测试文件包含以下蜂鸣片段:
- 1500 Hz 持续100 ms
- 1900 Hz 持续300 ms
- 1200 Hz 持续10 ms
- 1900 Hz 持续300 ms
现有方案会遗漏所有短于100ms的蜂鸣,还存在频率检测偏高的问题,原代码如下:
import numpy as np from scipy.io import wavfile def extract_frequencies_and_durations(wav_path, frame_duration=0.02, freq_tolerance=5): rate, audio_data = wavfile.read(wav_path) if audio_data.ndim > 1: audio_data = np.mean(audio_data, axis=1) frame_length = int(rate * frame_duration) frequencies = [] durations = [] current_freq = None current_duration = 0 for i in range(0, len(audio_data), frame_length): frame = audio_data[i:i + frame_length] fft_spectrum = np.fft.fft(frame) freqs = np.fft.fftfreq(len(fft_spectrum), 1 / rate) magnitude = np.abs(fft_spectrum) dominant_freq = abs(freqs[np.argmax(magnitude)]) if current_freq is None or abs(dominant_freq - current_freq) > freq_tolerance: if current_freq is not None: frequencies.append(current_freq) durations.append(current_duration) current_freq = dominant_freq current_duration = frame_duration else: current_duration += frame_duration if current_freq is not None: frequencies.append(current_freq) durations.append(current_duration) return frequencies, durations # output wav_path = 'test.wav' frequencies, durations = extract_frequencies_and_durations(wav_path) for freq, dur in zip(frequencies, durations): print(f"{freq:.2f}Hz - {dur * 1000:.2f}ms")
原代码核心问题
- 非重叠固定帧长:20ms帧长且步进等于帧长,10ms的短蜂鸣仅覆盖部分帧,易被前后帧信号淹没,无法被识别为独立段。
- 频谱泄漏与分辨率不足:未加窗导致FFT频谱泄漏,直接取峰值bin的频率会产生偏移;固定帧长的频率分辨率(采样率/帧长)偏低,无法精准定位真实频率。
- 简单频率判定逻辑:仅通过单帧频率与当前频率的差值判断切换,短帧信号易受相邻帧干扰,导致漏检。
修复后的代码
import numpy as np from scipy.io import wavfile def extract_frequencies_and_durations(wav_path, frame_duration=0.005, hop_duration=0.001, freq_tolerance=10, min_segment_duration=0.005): # 读取音频并转单声道 rate, audio_data = wavfile.read(wav_path) if audio_data.ndim > 1: audio_data = np.mean(audio_data, axis=1).astype(np.int16) # 归一化处理 audio_data = audio_data / np.iinfo(np.int16).max frame_length = int(rate * frame_duration) hop_length = int(rate * hop_duration) # 预加汉宁窗减少频谱泄漏 window = np.hanning(frame_length) frequencies = [] durations = [] current_freq = None current_duration = 0 current_segments = [] # 遍历重叠帧 for i in range(0, len(audio_data) - frame_length + 1, hop_length): frame = audio_data[i:i+frame_length] * window # 计算FFT并提取正频率部分 fft_spectrum = np.fft.fft(frame)[:frame_length//2] freqs = np.fft.fftfreq(frame_length, 1/rate)[:frame_length//2] magnitude = np.abs(fft_spectrum) # 抛物线插值提升频率精度 peak_idx = np.argmax(magnitude) if 0 < peak_idx < len(magnitude)-1: left, mid, right = magnitude[peak_idx-1], magnitude[peak_idx], magnitude[peak_idx+1] delta = (right - left) / (2 * (2*mid - left - right)) dominant_freq = freqs[peak_idx] + delta * (freqs[1] - freqs[0]) else: dominant_freq = freqs[peak_idx] # 基于平均频率的切换判定 if current_freq is None: current_freq = dominant_freq current_duration = hop_duration current_segments.append(dominant_freq) else: avg_current = np.mean(current_segments) if abs(dominant_freq - avg_current) > freq_tolerance: # 保存有效段(过滤过短噪声) if current_duration >= min_segment_duration: frequencies.append(avg_current) durations.append(current_duration) # 切换到新频率段 current_freq = dominant_freq current_duration = hop_duration current_segments = [dominant_freq] else: current_duration += hop_duration current_segments.append(dominant_freq) # 处理最后一段 if current_freq is not None and current_duration >= min_segment_duration: frequencies.append(np.mean(current_segments)) durations.append(current_duration) return frequencies, durations # 测试执行 wav_path = 'test.wav' frequencies, durations = extract_frequencies_and_durations(wav_path) for freq, dur in zip(frequencies, durations): print(f"{round(freq)}Hz - {dur * 1000:.1f}ms")
关键改进点
- 重叠短帧:使用5ms帧长+1ms步进的重叠帧,确保10ms短蜂鸣能被多帧捕获,避免漏检。
- 频谱优化:汉宁窗减少泄漏,抛物线插值提升频率检测精度,解决频率偏高问题。
- 稳定判定逻辑:基于当前段的平均频率判断切换,避免单帧噪声干扰,提升短信号识别稳定性。
- 灵活阈值:可自定义最小检测时长,过滤无效噪声同时保留有效短信号。
内容的提问来源于stack exchange,提问作者kotek
相关产品推荐
相关产品推荐

