如何借助Azure认知服务实现可打断的同步听与语音合成
实现Azure认知服务语音合成与识别的打断功能问题
我正在基于Azure认知服务开发一款支持语音输出与语音识别同步进行的聊天机器人,目前已实现两个核心函数:
已实现的核心函数
语音合成函数
def speak(input,voice="en-US-ChristopherNeural"): audio_config = speechsdk.audio.AudioOutputConfig(use_default_speaker=True) speech_config.speech_synthesis_voice_name=voice speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config) speech_synthesis_result = speech_synthesizer.speak_text_async(input).get() if speech_synthesis_result.reason == speechsdk.ResultReason.Canceled: cancellation_details = speech_synthesis_result.cancellation_details print("Azure Speech synthesis canceled: {}".format(cancellation_details.reason)) return True
语音识别函数
def listen(language): speech_config.speech_recognition_language=language audio_config = speechsdk.audio.AudioConfig(use_default_microphone=True) speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config) print("Speak into your microphone.") speech_recognition_result = speech_recognizer.recognize_once_async().get() if speech_recognition_result.reason == speechsdk.ResultReason.RecognizedSpeech: print("Recognized: {}".format(speech_recognition_result.text)) return speech_recognition_result.text
需求与问题
我需要实现用户说话时可打断正在进行的语音输出的功能,这要求持续监听麦克风输入以检测用户的语音触发。
尝试过多线程方案,但被speech_synthesis_result = speech_synthesizer.speak_text_async(input).get()这行代码阻塞。之后编写了异步代码,但程序无法正常发声,请求解决。
问题代码
import asyncio import os import azure.cognitiveservices.speech as speechsdk from azure.cognitiveservices.speech import SpeechConfig, SpeechSynthesisOutputFormat, SpeechSynthesizer # Replace with your own subscription key and region identifier speech_config = speechsdk.SpeechConfig(subscription=os.environ.get('SPEECH_KEY'), region=os.environ.get('SPEECH_REGION')) # Define the phrase to be spoken phrase = "Hello, I'm a chatbot. How can I help you today? You can interrupt me whenever you want" async def listen_for_user_input(): speech_config.speech_recognition_language="en-US-ChristopherNeural" audio_config = speechsdk.audio.AudioConfig(use_default_microphone=True) speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config) result = await speech_recognizer.start_continuous_recognition_async() if result.reason == speechsdk.ResultReason.RecognizedSpeech: print("Recognized: {}".format(result.text)) await speech_recognizer.stop_continuous_recognition_async() pass async def speak_phrase(phrase): audio_config = speechsdk.audio.AudioOutputConfig(use_default_speaker=True) synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config) # Speak the defined phrase result = await synthesizer.speak_text_async(phrase) if result.reason == speechsdk.ResultReason.Canceled: cancellation_details = result.cancellation_details print("Azure Speech synthesis canceled: {}".format(cancellation_details.reason)) # Start the program async def main(): task1 = asyncio.create_task(speak_phrase(phrase)) task2 = asyncio.create_task(listen_for_user_input()) done, pending = await asyncio.wait([task1, task2], return_when=asyncio.FIRST_COMPLETED) for task in pending: task.cancel()
解决方案
问题分析
- 异步适配错误:Azure Speech SDK的
async方法返回的是Future对象,不是原生asyncio可直接await的协程,需要用asyncio.wrap_future转换。 - 连续识别逻辑错误:
start_continuous_recognition_async只是启动识别,不会直接返回识别结果,需要通过注册回调函数获取。 - 合成中断逻辑缺失:没有在检测到用户语音时主动停止正在进行的语音合成。
修正后的代码
import asyncio import os import azure.cognitiveservices.speech as speechsdk # 配置Azure Speech服务 speech_config = speechsdk.SpeechConfig(subscription=os.environ.get('SPEECH_KEY'), region=os.environ.get('SPEECH_REGION')) phrase = "Hello, I'm a chatbot. How can I help you today? You can interrupt me whenever you want" # 全局变量用于控制合成中断 interrupt_triggered = False def handle_recognized_result(evt): global interrupt_triggered if evt.result.reason == speechsdk.ResultReason.RecognizedSpeech: print(f"Recognized: {evt.result.text}") interrupt_triggered = True async def listen_for_user_input(): audio_config = speechsdk.audio.AudioConfig(use_default_microphone=True) speech_recognizer = speechsdk.SpeechRecognizer( speech_config=speech_config, audio_config=audio_config, speech_recognition_language="en-US" ) # 注册识别结果回调 speech_recognizer.recognized.connect(handle_recognized_result) # 启动连续识别 await asyncio.wrap_future(speech_recognizer.start_continuous_recognition_async()) # 等待中断触发 while not interrupt_triggered: await asyncio.sleep(0.1) # 停止识别 await asyncio.wrap_future(speech_recognizer.stop_continuous_recognition_async()) async def speak_phrase(phrase): global interrupt_triggered audio_config = speechsdk.audio.AudioOutputConfig(use_default_speaker=True) synthesizer = speechsdk.SpeechSynthesizer( speech_config=speech_config, audio_config=audio_config ) # 启动合成任务 synthesis_future = synthesizer.speak_text_async(phrase) # 轮询检查是否需要中断 while not synthesis_future.done() and not interrupt_triggered: await asyncio.sleep(0.1) # 如果触发中断,停止合成 if interrupt_triggered: synthesizer.stop_speaking_async() print("Speech synthesis interrupted by user input") result = await asyncio.wrap_future(synthesis_future) if result.reason == speechsdk.ResultReason.Canceled: cancellation_details = result.cancellation_details print(f"Azure Speech synthesis canceled: {cancellation_details.reason}") async def main(): global interrupt_triggered interrupt_triggered = False task1 = asyncio.create_task(speak_phrase(phrase)) task2 = asyncio.create_task(listen_for_user_input()) done, pending = await asyncio.wait([task1, task2], return_when=asyncio.FIRST_COMPLETED) for task in pending: task.cancel() try: await task except asyncio.CancelledError: pass if __name__ == "__main__": asyncio.run(main())
关键改进点
- 使用
asyncio.wrap_future将SDK的Future对象转换为asyncio兼容的协程。 - 通过回调函数
handle_recognized_result实时获取语音识别结果,触发中断标记。 - 在语音合成任务中轮询中断标记,一旦触发就调用
stop_speaking_async停止合成。 - 修复了连续语音识别的逻辑,确保持续监听用户输入。
内容的提问来源于stack exchange,提问作者Fabian
相关产品推荐
相关产品推荐

