Python语音助手静音检测异常:is_silence无语音时返回False问题
Python语音助手静音检测异常排查与解决
问题背景
正在开发小型Python语音助手,基于Vosk实现语音识别、PyAudio录制音频、Numpy计算音频能量、Wave保存音频文件。当前在实现静音检测判断用户语句结束时遇到异常:即使不说话(仅存在轻微背景噪音),is_silence方法仍返回False。尝试过更换pydub、webrtcvad等工具,以及调整阈值参数,均未解决问题。运行环境为MacOS + Python3.12虚拟环境。
原始代码片段
import vosk import json import pyaudio import numpy as np import wave def is_silence(audio_chunk, silence_thresh=-50, energy_thresh=0.001): # Convert the audio chunk to a NumPy array of 16-bit integers audio_chunk = np.frombuffer(audio_chunk, dtype=np.int16) # Calculate the energy of the audio chunk energy = np.sum(audio_chunk.astype(np.float32) ** 2) / float(len(audio_chunk)) # Check if both energy and maximum absolute value are below thresholds return energy < energy_thresh and np.max(np.abs(audio_chunk)) < silence_thresh def record_audio(output_file, max_silence_duration=1500, sample_rate=16000, chunk_size=1024, silence_thresh=-50): # Initialize PyAudio p = pyaudio.PyAudio() # Open a stream for recording audio stream = p.open(format=pyaudio.paInt16, channels=1, rate=sample_rate, input=True, frames_per_buffer=chunk_size) # Initialize variables for storing audio frames and counting consecutive silence frames = [] consecutive_silence_counter = 0 print("Recording...") while True: # Read a chunk of audio data from the stream data = stream.read(chunk_size) frames.append(data) # Convert the audio chunk to a NumPy array audio_chunk = np.frombuffer(data, dtype=np.int16) # Check if the audio chunk is silence if is_silence(audio_chunk, silence_thresh=silence_thresh): consecutive_silence_counter += 1 else: consecutive_silence_counter = 0 # Break the loop if consecutive silence reaches the specified duration if consecutive_silence_counter * chunk_size >= (max_silence_duration // chunk_size): break print("Recording finished.") # Stop and close the audio stream, and terminate PyAudio stream.stop_stream() stream.close() p.terminate() # Write the recorded frames to a WAV file with wave.open(output_file, 'wb') as wf: wf.setnchannels(1) wf.setsampwidth(p.get_sample_size(pyaudio.paInt16)) wf.setframerate(sample_rate) wf.writeframes(b''.join(frames)) # recognize_speech function if __name__ == "__main__": # Path to the Vosk model model_path = "path/to/vosk-model-en-us-0.22-lgraph" # Path to the output audio file output_file_path = "path to audio file" # Record audio and save it to the specified output file audio_data = record_audio(output_file=output_file_path) # Recognize speech from the recorded audio using the Vosk model recognized_text = recognize_speech(model_path, audio_data) # Print the recognized text print("Recognized text:", recognized_text)
核心问题排查与解决方案
1. 修复阈值判断逻辑错误
原is_silence函数存在致命逻辑错误:
- 16位音频的取值范围是
-32768到32767,np.max(np.abs(audio_chunk))得到的是非负的绝对值最大值,但你设置的silence_thresh=-50是负数,导致这个条件永远为False,最终整个函数返回False。
修改后的is_silence函数:
def is_silence(audio_chunk, silence_thresh=50, energy_thresh=0.001): audio_chunk = np.frombuffer(audio_chunk, dtype=np.int16) # 计算音频能量 energy = np.sum(audio_chunk.astype(np.float32) ** 2) / float(len(audio_chunk)) # 获取音频块的最大绝对值 max_abs_value = np.max(np.abs(audio_chunk)) # 调试输出(可选,用于观察实际数值调整阈值) print(f"当前能量: {energy:.6f}, 最大绝对值: {max_abs_value}") # 修正逻辑:判断绝对值最大值小于正数阈值,同时能量低于阈值 return energy < energy_thresh and max_abs_value < silence_thresh
2. 修正录音停止条件的计算错误
原record_audio函数中,停止录制的条件单位不匹配:
consecutive_silence_counter是音频块的数量,chunk_size是每个块的采样数,两者相乘无法直接和毫秒级的max_silence_duration比较。
修改后的停止判断逻辑:
def record_audio(output_file, max_silence_duration=1500, sample_rate=16000, chunk_size=1024, silence_thresh=50): p = pyaudio.PyAudio() stream = p.open(format=pyaudio.paInt16, channels=1, rate=sample_rate, input=True, frames_per_buffer=chunk_size) frames = [] consecutive_silence_counter = 0 # 计算单个音频块的时长(毫秒) chunk_duration_ms = (chunk_size / sample_rate) * 1000 print("Recording...") while True: data = stream.read(chunk_size) frames.append(data) if is_silence(data, silence_thresh=silence_thresh): consecutive_silence_counter += 1 else: consecutive_silence_counter = 0 # 修正:连续静音总时长 >= 设置的最大静音时长时停止 if consecutive_silence_counter * chunk_duration_ms >= max_silence_duration: break print("Recording finished.") stream.stop_stream() stream.close() p.terminate() with wave.open(output_file, 'wb') as wf: wf.setnchannels(1) wf.setsampwidth(p.get_sample_size(pyaudio.paInt16)) wf.setframerate(sample_rate) wf.writeframes(b''.join(frames))
3. 适配MacOS音频特性
MacOS默认音频输入增益可能过高,导致背景噪音被放大。可通过两种方式优化:
- 手动调整系统音频输入增益:打开「系统设置」→「声音」→「输入」,降低麦克风增益。
- 代码中添加音频归一化处理(可选):在
is_silence函数中加入:# 归一化音频(避免增益过高导致数值异常) if np.max(np.abs(audio_chunk)) != 0: audio_chunk = audio_chunk / np.max(np.abs(audio_chunk))
4. 补充完整语音识别函数
原代码中recognize_speech函数未实现,补充完整逻辑:
def recognize_speech(model_path, audio_file_path): model = vosk.Model(model_path) rec = vosk.KaldiRecognizer(model, 16000) with wave.open(audio_file_path, "rb") as wf: while True: data = wf.readframes(4000) if len(data) == 0: break if rec.AcceptWaveform(data): result = json.loads(rec.Result()) return result.get('text', '') # 获取最终识别结果 final_result = json.loads(rec.FinalResult()) return final_result.get('text', '')
完整修正后的代码
整合所有修改后的完整代码可直接运行测试,根据调试输出的能量和绝对值数值,微调silence_thresh和energy_thresh参数即可适配你的环境。
内容的提问来源于stack exchange,提问作者Dima Paharski
相关产品推荐
相关产品推荐

