You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DeepSpeech语音识别脚本无转录输出问题排查求助

问题:DeepSpeech无法转录麦克风输入的语音

已下载所有依赖包并安装DeepSpeech模型,代码无报错,但运行后无法将麦克风输入的语音转录为文本。

代码示例

import deepspeech
import numpy as np
import pyaudio
import wave

# Set the path to the DeepSpeech model and scorer
MODEL_PATH = 'E:\Pie-Infocomm\Deep-speech\deepspeech-0.9.3-models.pbmm'
SCORER_PATH = 'E:\Pie-Infocomm\Deep-speech\deepspeech-0.9.3-models.scorer'

def load_model():
    model = deepspeech.Model(MODEL_PATH)
    model.enableExternalScorer(SCORER_PATH)
    return model

def transcribe_audio(model, audio_data):
    return model.stt(audio_data)

def main():
    model = load_model()

    CHUNK = 1024
    FORMAT = pyaudio.paInt16
    CHANNELS = 1
    RATE = 16000

    p = pyaudio.PyAudio()

    stream = p.open(format=FORMAT,
                    channels=CHANNELS,
                    rate=RATE,
                    input=True,
                    frames_per_buffer=CHUNK)

    print("Listening...")

    while True:
        try:
            audio_data = stream.read(CHUNK)
            audio_array = np.frombuffer(audio_data, dtype=np.int16)
            text = transcribe_audio(model, audio_array)
            print("Text:", text)
        except KeyboardInterrupt:
            break

    print("Finished recording")

    stream.stop_stream()
    stream.close()
    p.terminate()

if __name__ == "__main__":
    main()

运行输出

[Running] python -u "e:\Pie-Infocomm\Deep-speech\Test-Deepspeech.py"
TensorFlow: v2.3.0-6-g23ad988fcd
DeepSpeech: v0.9.3-0-gf2e9c858
Listening...
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 
Text: 

问题原因及解决方案

1. 音频片段过短

当前代码每次仅传入1024个采样点(约64毫秒)的音频,DeepSpeech需要足够长的语音片段才能有效识别,这么短的输入无法触发准确的转录。

2. 未使用流式识别API

DeepSpeech提供了专门的流式识别接口,适合处理实时麦克风输入,而非单次调用stt()处理小块音频。

3. 路径转义(潜在问题)

Windows路径中的反斜杠需要转义,或使用原始字符串避免解析错误,比如将路径改为:

MODEL_PATH = r'E:\Pie-Infocomm\Deep-speech\deepspeech-0.9.3-models.pbmm'
SCORER_PATH = r'E:\Pie-Infocomm\Deep-speech\deepspeech-0.9.3-models.scorer'

修改后的代码

import deepspeech
import numpy as np
import pyaudio

# 使用原始字符串避免路径转义问题
MODEL_PATH = r'E:\Pie-Infocomm\Deep-speech\deepspeech-0.9.3-models.pbmm'
SCORER_PATH = r'E:\Pie-Infocomm\Deep-speech\deepspeech-0.9.3-models.scorer'

def load_model():
    model = deepspeech.Model(MODEL_PATH)
    model.enableExternalScorer(SCORER_PATH)
    return model

def main():
    model = load_model()

    CHUNK = 1024
    FORMAT = pyaudio.paInt16
    CHANNELS = 1
    RATE = 16000

    p = pyaudio.PyAudio()

    stream = p.open(format=FORMAT,
                    channels=CHANNELS,
                    rate=RATE,
                    input=True,
                    frames_per_buffer=CHUNK)

    print("Listening... Press Ctrl+C to stop")

    # 创建流式识别对象
    ds_stream = model.createStream()

    try:
        while True:
            audio_data = stream.read(CHUNK)
            audio_array = np.frombuffer(audio_data, dtype=np.int16)
            # 向流式对象喂入音频数据
            ds_stream.feedAudioContent(audio_array)
            # 实时获取当前转录结果
            text = ds_stream.intermediateDecode()
            # 清空当前行并打印最新结果(优化输出体验)
            print(f"\rCurrent Text: {text}", end="")
    except KeyboardInterrupt:
        # 结束流式识别,获取最终结果
        final_text = ds_stream.finishStream()
        print(f"\nFinal Transcription: {final_text}")

    print("\nFinished recording")
    stream.stop_stream()
    stream.close()
    p.terminate()

if __name__ == "__main__":
    main()

修改说明

  • 使用createStream()、feedAudioContent()和intermediateDecode()实现流式实时识别,积累足够音频后输出结果。
  • 优化输出格式,实时更新当前转录文本,避免重复空行。
  • 修复Windows路径的转义问题,使用原始字符串保证路径正确。

内容的提问来源于stack exchange,提问作者Anjani Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 12:32:23