使用Python Sphinx转录实时音频时遇UnicodeDecodeError求助
问题分析
你的错误根源有这几点:
LiveSpeech是pocketsphinx封装的麦克风实时语音识别工具,它默认从系统麦克风读取音频,你直接传入网络音频二进制数据时,内部逻辑错误地将二进制音频尝试按UTF-8解码,触发UnicodeDecodeError。- 代码中开启了PyAudio的输入流但完全未使用,属于冗余代码,直接删除即可。
- 未验证网络音频流是否是pocketsphinx要求的16kHz、单声道、16位无符号PCM格式,如果是带容器的音频(如WAV、MP3),直接喂给识别模块会导致格式不兼容。
修正方案
情况1:网络流是原始16kHz单声道PCM数据
改用pocketsphinx.Decoder手动处理音频流,代码如下:
import requests from pocketsphinx import Decoder def speech_to_text(audio_stream): # 配置Decoder参数,匹配音频格式 config = Decoder.default_config() config.set_string('-hmm', 'en-us') # 替换为你的模型路径,默认自带英文模型 config.set_string('-lm', 'en-us.lm.bin') config.set_string('-dict', 'cmudict-en-us.dict') config.set_float('-samprate', 16000) config.set_int('-nfft', 2048) decoder = Decoder(config) decoder.start_utt() # 开始识别会话 for chunk in audio_stream.iter_content(chunk_size=1024): if chunk: decoder.process_raw(chunk, False, False) # 处理原始PCM数据 # 检查是否识别到完整短语 if decoder.hyp() is not None: yield decoder.hyp().hypstr decoder.end_utt() decoder.start_utt() def main(): url = "AUDIO_URL" # 明确获取二进制数据,避免requests自动尝试解码文本 audio_stream = requests.get(url, stream=True) audio_stream.raise_for_status() # 检查请求是否成功 print("Listening to audio stream...") try: for phrase in speech_to_text(audio_stream): print("Recognized text:", phrase) except KeyboardInterrupt: print("Stopped listening.") if __name__ == "__main__": main()
情况2:网络流是压缩格式(如MP3、WAV)
需要先将音频解码为原始PCM,推荐用pydub库处理:
- 先安装依赖:
pip install pydub
- 修正后的代码:
import requests from pocketsphinx import Decoder from pydub import AudioSegment from io import BytesIO def speech_to_text(audio_stream): # 将音频流转换为AudioSegment,自动识别格式 audio = AudioSegment.from_file(BytesIO(audio_stream.content)) # 转换为pocketsphinx要求的格式:16kHz、单声道、16位PCM audio = audio.set_frame_rate(16000).set_channels(1).set_sample_width(2) raw_pcm = audio.raw_data # 配置Decoder config = Decoder.default_config() config.set_string('-hmm', 'en-us') config.set_string('-lm', 'en-us.lm.bin') config.set_string('-dict', 'cmudict-en-us.dict') config.set_float('-samprate', 16000) config.set_int('-nfft', 2048) decoder = Decoder(config) decoder.start_utt() decoder.process_raw(raw_pcm, False, True) # 处理完整PCM数据 if decoder.hyp() is not None: yield decoder.hyp().hypstr decoder.end_utt() def main(): url = "AUDIO_URL" audio_stream = requests.get(url) audio_stream.raise_for_status() print("Processing audio stream...") for phrase in speech_to_text(audio_stream): print("Recognized text:", phrase) if __name__ == "__main__": main()
关键注意事项
- 确保pocketsphinx的模型文件(hmm、lm、dict)路径正确,默认库自带英文模型,使用其他语言需替换对应模型。
- 如果是流式传输的压缩音频,需要用
ffmpeg结合管道做流式解码,上述pydub方式适合非流式的完整音频文件。 - 已移除无用的PyAudio代码,因为你不需要从麦克风捕获音频,仅需处理网络流。
内容的提问来源于stack exchange,提问作者Abraham Arnold
相关产品推荐
相关产品推荐

