You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Sphinx转录实时音频时遇UnicodeDecodeError求助

问题分析

你的错误根源有这几点:

  • LiveSpeech是pocketsphinx封装的麦克风实时语音识别工具,它默认从系统麦克风读取音频,你直接传入网络音频二进制数据时,内部逻辑错误地将二进制音频尝试按UTF-8解码,触发UnicodeDecodeError。
  • 代码中开启了PyAudio的输入流但完全未使用,属于冗余代码,直接删除即可。
  • 未验证网络音频流是否是pocketsphinx要求的16kHz、单声道、16位无符号PCM格式,如果是带容器的音频(如WAV、MP3),直接喂给识别模块会导致格式不兼容。
修正方案

情况1:网络流是原始16kHz单声道PCM数据

改用pocketsphinx.Decoder手动处理音频流,代码如下:

import requests
from pocketsphinx import Decoder

def speech_to_text(audio_stream):
    # 配置Decoder参数,匹配音频格式
    config = Decoder.default_config()
    config.set_string('-hmm', 'en-us')  # 替换为你的模型路径,默认自带英文模型
    config.set_string('-lm', 'en-us.lm.bin')
    config.set_string('-dict', 'cmudict-en-us.dict')
    config.set_float('-samprate', 16000)
    config.set_int('-nfft', 2048)
    
    decoder = Decoder(config)
    decoder.start_utt()  # 开始识别会话
    
    for chunk in audio_stream.iter_content(chunk_size=1024):
        if chunk:
            decoder.process_raw(chunk, False, False)  # 处理原始PCM数据
            # 检查是否识别到完整短语
            if decoder.hyp() is not None:
                yield decoder.hyp().hypstr
                decoder.end_utt()
                decoder.start_utt()

def main():
    url = "AUDIO_URL"
    # 明确获取二进制数据,避免requests自动尝试解码文本
    audio_stream = requests.get(url, stream=True)
    audio_stream.raise_for_status()  # 检查请求是否成功

    print("Listening to audio stream...")

    try:
        for phrase in speech_to_text(audio_stream):
            print("Recognized text:", phrase)
    except KeyboardInterrupt:
        print("Stopped listening.")

if __name__ == "__main__":
    main()

情况2:网络流是压缩格式(如MP3、WAV)

需要先将音频解码为原始PCM,推荐用pydub库处理:

  1. 先安装依赖:
pip install pydub
  1. 修正后的代码:
import requests
from pocketsphinx import Decoder
from pydub import AudioSegment
from io import BytesIO

def speech_to_text(audio_stream):
    # 将音频流转换为AudioSegment,自动识别格式
    audio = AudioSegment.from_file(BytesIO(audio_stream.content))
    # 转换为pocketsphinx要求的格式:16kHz、单声道、16位PCM
    audio = audio.set_frame_rate(16000).set_channels(1).set_sample_width(2)
    raw_pcm = audio.raw_data

    # 配置Decoder
    config = Decoder.default_config()
    config.set_string('-hmm', 'en-us')
    config.set_string('-lm', 'en-us.lm.bin')
    config.set_string('-dict', 'cmudict-en-us.dict')
    config.set_float('-samprate', 16000)
    config.set_int('-nfft', 2048)
    
    decoder = Decoder(config)
    decoder.start_utt()
    decoder.process_raw(raw_pcm, False, True)  # 处理完整PCM数据
    
    if decoder.hyp() is not None:
        yield decoder.hyp().hypstr
    decoder.end_utt()

def main():
    url = "AUDIO_URL"
    audio_stream = requests.get(url)
    audio_stream.raise_for_status()

    print("Processing audio stream...")
    for phrase in speech_to_text(audio_stream):
        print("Recognized text:", phrase)

if __name__ == "__main__":
    main()
关键注意事项
  • 确保pocketsphinx的模型文件(hmm、lm、dict)路径正确,默认库自带英文模型,使用其他语言需替换对应模型。
  • 如果是流式传输的压缩音频,需要用ffmpeg结合管道做流式解码,上述pydub方式适合非流式的完整音频文件。
  • 已移除无用的PyAudio代码,因为你不需要从麦克风捕获音频,仅需处理网络流。

内容的提问来源于stack exchange,提问作者Abraham Arnold

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 23:35:24