You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Cognitive Services语音转写异常:生成的TXT文件为空

音频转写TXT为空:recognized回调未触发的解决办法

尝试将音频文件转写内容保存至TXT文件,但生成的文件为空。控制台仅输出RECOGNIZING: hello,未触发recognized回调写入文件,实现代码如下:

speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)
audio_config = speechsdk.audio.AudioConfig(filename=audio_file)

speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config)
done = False

# Define a callback function to handle the result
def handle_result(result):
    print("Got result:", result.text)
    with open("transcript.txt", "a") as f:
        f.write(result.text + "\n")

# Start the continuous transcription process
speech_recognizer.recognizing.connect(lambda evt: print('RECOGNIZING: {}'.format(evt.result.text)))
speech_recognizer.recognized.connect(handle_result)
speech_recognizer.session_stopped.connect(lambda evt: print('SESSION STOPPED {}'.format(evt)))

speech_recognizer.start_continuous_recognition()
time.sleep(15) 
speech_recognizer.stop_continuous_recognition()

解决步骤

  • 替换固定sleep为事件驱动等待
    固定time.sleep(15)会强制终止未完成的识别过程,导致recognized回调无法触发。改用会话结束事件控制停止时机:

    done = False
    
    def stop_recognition(evt):
        print('SESSION STOPPED {}'.format(evt))
        speech_recognizer.stop_continuous_recognition()
        nonlocal done
        done = True
    
    # 绑定会话停止和取消事件
    speech_recognizer.session_stopped.connect(stop_recognition)
    speech_recognizer.canceled.connect(stop_recognition)
    
    speech_recognizer.start_continuous_recognition()
    # 等待识别完成
    while not done:
        time.sleep(0.5)
    
  • 添加错误捕获回调排查问题
    未监听canceled回调会错过识别过程中的错误信息,添加后可快速定位授权、音频格式或网络问题:

    def handle_canceled(evt):
        print('CANCELED: {}'.format(evt.reason))
        if evt.reason == speechsdk.CancellationReason.Error:
            print('ERROR details: {}'.format(evt.error_details))
    
    speech_recognizer.canceled.connect(handle_canceled)
    
  • 确保音频触发最终识别结果
    recognized回调对应稳定的最终识别结果,若音频过短或无结尾静音,SDK可能无法确认识别完成。可以给音频末尾添加2-3秒静音,或检查音频文件是否完整。

  • 验证音频格式兼容性
    确认音频文件为Speech SDK支持的格式(如WAV、MP3),若格式不兼容会导致识别中断。建议转换为16kHz、16位、单声道的PCM WAV格式测试。

内容的提问来源于stack exchange,提问作者Genaro Raúl Mateu Lagunes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 04:45:22