如何用Python不存本地直接将音频响应传入Azure语音识别?PullAudioStream异常
解决方案:直接流式音频传入Azure语音识别(修复PullAudioStream异常)
1. 核心问题定位
你遇到的PullAudioStream异常,大多是因为未正确实现抽象类的必填方法,或是混淆了流操作与识别回调的绑定逻辑。get_property默认返回空值,若要自定义属性必须显式重写;而识别回调是绑定在语音识别器实例上,而非流对象本身。
2. 正确实现自定义PullAudioStream
继承PullAudioInputStream时,必须实现read()方法来提供音频数据,若需要自定义属性则重写get_property()。以下是基础模板:
import azure.cognitiveservices.speech as speechsdk from io import BytesIO class CustomPullAudioStream(speechsdk.audio.PullAudioInputStream): def __init__(self, audio_source: BytesIO): super().__init__() self.audio_source = audio_source # 必须实现的核心方法:返回音频字节,结束时返回空字节串 def read(self, buffer: memoryview) -> int: chunk = self.audio_source.read(len(buffer)) if not chunk: return 0 buffer[:len(chunk)] = chunk return len(chunk) # 重写get_property以支持自定义属性(按需实现) def get_property(self, id: speechsdk.PropertyId) -> str: # 示例:返回音频格式属性 if id == speechsdk.PropertyId.SpeechServiceConnection_AudioFormat: return "audio/wav; codecs=audio/pcm; samplerate=16000" # 其他属性返回默认值 return super().get_property(id)
3. 绑定识别回调并启动识别
回调逻辑需要绑定到语音识别器的recognizing(中间结果)和recognized(最终结果)事件,而非流对象。完整流程示例:
def main(): # 1. 配置Azure语音服务 speech_config = speechsdk.SpeechConfig(subscription="你的订阅密钥", region="你的区域") speech_config.speech_recognition_language = "zh-CN" # 2. 模拟从网络/内存获取音频(不保存到本地) # 示例:这里用BytesIO模拟音频源,实际可替换为网络请求的响应流 import requests audio_url = "你的音频文件URL" audio_response = requests.get(audio_url, stream=True) audio_stream = BytesIO(audio_response.content) # 3. 创建自定义拉流对象 custom_audio_input = CustomPullAudioStream(audio_stream) audio_config = speechsdk.audio.AudioConfig(stream=custom_audio_input) # 4. 绑定回调函数 def recognizing_handler(evt): print(f"中间识别结果: {evt.result.text}") def recognized_handler(evt): if evt.result.reason == speechsdk.ResultReason.RecognizedSpeech: print(f"最终识别结果: {evt.result.text}") elif evt.result.reason == speechsdk.ResultReason.NoMatch: print(f"无匹配结果: {evt.result.no_match_details}") # 5. 创建识别器并启动识别 speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config) speech_recognizer.recognizing.connect(recognizing_handler) speech_recognizer.recognized.connect(recognized_handler) # 启动连续识别(或用recognize_once()单次识别) speech_recognizer.start_continuous_recognition() # 等待识别完成(实际场景可根据业务逻辑处理) import time time.sleep(10) speech_recognizer.stop_continuous_recognition() if __name__ == "__main__": main()
关键注意事项
read()方法必须严格遵循:每次读取不超过buffer长度的字节,音频结束时返回0,否则会导致流异常。- 若不需要自定义属性,可省略
get_property的重写,直接使用父类默认实现。 - 回调函数的参数是
SpeechRecognitionEventArgs子类,需通过evt.result获取识别内容。 - 音频格式必须与Azure服务要求匹配(默认支持16kHz、16位、单声道PCM),若格式不同需在
get_property中指定正确的SpeechServiceConnection_AudioFormat值。
内容的提问来源于stack exchange,提问作者Vishwa Karthi
相关产品推荐
相关产品推荐

