You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python不存本地直接将音频响应传入Azure语音识别?PullAudioStream异常

解决方案:直接流式音频传入Azure语音识别(修复PullAudioStream异常)

1. 核心问题定位

你遇到的PullAudioStream异常,大多是因为未正确实现抽象类的必填方法,或是混淆了流操作与识别回调的绑定逻辑。get_property默认返回空值,若要自定义属性必须显式重写;而识别回调是绑定在语音识别器实例上,而非流对象本身。

2. 正确实现自定义PullAudioStream

继承PullAudioInputStream时,必须实现read()方法来提供音频数据,若需要自定义属性则重写get_property()。以下是基础模板:

import azure.cognitiveservices.speech as speechsdk
from io import BytesIO

class CustomPullAudioStream(speechsdk.audio.PullAudioInputStream):
    def __init__(self, audio_source: BytesIO):
        super().__init__()
        self.audio_source = audio_source

    # 必须实现的核心方法:返回音频字节,结束时返回空字节串
    def read(self, buffer: memoryview) -> int:
        chunk = self.audio_source.read(len(buffer))
        if not chunk:
            return 0
        buffer[:len(chunk)] = chunk
        return len(chunk)

    # 重写get_property以支持自定义属性(按需实现)
    def get_property(self, id: speechsdk.PropertyId) -> str:
        # 示例:返回音频格式属性
        if id == speechsdk.PropertyId.SpeechServiceConnection_AudioFormat:
            return "audio/wav; codecs=audio/pcm; samplerate=16000"
        # 其他属性返回默认值
        return super().get_property(id)

3. 绑定识别回调并启动识别

回调逻辑需要绑定到语音识别器的recognizing(中间结果)和recognized(最终结果)事件,而非流对象。完整流程示例:

def main():
    # 1. 配置Azure语音服务
    speech_config = speechsdk.SpeechConfig(subscription="你的订阅密钥", region="你的区域")
    speech_config.speech_recognition_language = "zh-CN"

    # 2. 模拟从网络/内存获取音频(不保存到本地)
    # 示例:这里用BytesIO模拟音频源,实际可替换为网络请求的响应流
    import requests
    audio_url = "你的音频文件URL"
    audio_response = requests.get(audio_url, stream=True)
    audio_stream = BytesIO(audio_response.content)

    # 3. 创建自定义拉流对象
    custom_audio_input = CustomPullAudioStream(audio_stream)
    audio_config = speechsdk.audio.AudioConfig(stream=custom_audio_input)

    # 4. 绑定回调函数
    def recognizing_handler(evt):
        print(f"中间识别结果: {evt.result.text}")

    def recognized_handler(evt):
        if evt.result.reason == speechsdk.ResultReason.RecognizedSpeech:
            print(f"最终识别结果: {evt.result.text}")
        elif evt.result.reason == speechsdk.ResultReason.NoMatch:
            print(f"无匹配结果: {evt.result.no_match_details}")

    # 5. 创建识别器并启动识别
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config)
    speech_recognizer.recognizing.connect(recognizing_handler)
    speech_recognizer.recognized.connect(recognized_handler)

    # 启动连续识别(或用recognize_once()单次识别)
    speech_recognizer.start_continuous_recognition()
    # 等待识别完成(实际场景可根据业务逻辑处理)
    import time
    time.sleep(10)
    speech_recognizer.stop_continuous_recognition()

if __name__ == "__main__":
    main()

关键注意事项

  • read()方法必须严格遵循:每次读取不超过buffer长度的字节,音频结束时返回0,否则会导致流异常。
  • 若不需要自定义属性,可省略get_property的重写,直接使用父类默认实现。
  • 回调函数的参数是SpeechRecognitionEventArgs子类,需通过evt.result获取识别内容。
  • 音频格式必须与Azure服务要求匹配(默认支持16kHz、16位、单声道PCM),若格式不同需在get_property中指定正确的SpeechServiceConnection_AudioFormat值。

内容的提问来源于stack exchange,提问作者Vishwa Karthi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 13:47:25