You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何借助Azure认知服务实现可打断的同步听与语音合成

实现Azure认知服务语音合成与识别的打断功能问题

我正在基于Azure认知服务开发一款支持语音输出与语音识别同步进行的聊天机器人,目前已实现两个核心函数:

已实现的核心函数

语音合成函数

def speak(input,voice="en-US-ChristopherNeural"):
    audio_config = speechsdk.audio.AudioOutputConfig(use_default_speaker=True)
    speech_config.speech_synthesis_voice_name=voice
    speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config)

    speech_synthesis_result = speech_synthesizer.speak_text_async(input).get()
    if speech_synthesis_result.reason == speechsdk.ResultReason.Canceled:
        cancellation_details = speech_synthesis_result.cancellation_details
        print("Azure Speech synthesis canceled: {}".format(cancellation_details.reason))

    return True

语音识别函数

def listen(language):
    speech_config.speech_recognition_language=language
    audio_config = speechsdk.audio.AudioConfig(use_default_microphone=True)
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config)

    print("Speak into your microphone.")
    speech_recognition_result = speech_recognizer.recognize_once_async().get()

    if speech_recognition_result.reason == speechsdk.ResultReason.RecognizedSpeech:
        print("Recognized: {}".format(speech_recognition_result.text))
        return speech_recognition_result.text

需求与问题

我需要实现用户说话时可打断正在进行的语音输出的功能,这要求持续监听麦克风输入以检测用户的语音触发。

尝试过多线程方案,但被speech_synthesis_result = speech_synthesizer.speak_text_async(input).get()这行代码阻塞。之后编写了异步代码,但程序无法正常发声,请求解决。

问题代码

import asyncio
import os
import azure.cognitiveservices.speech as speechsdk
from azure.cognitiveservices.speech import SpeechConfig, SpeechSynthesisOutputFormat, SpeechSynthesizer

# Replace with your own subscription key and region identifier
speech_config = speechsdk.SpeechConfig(subscription=os.environ.get('SPEECH_KEY'), region=os.environ.get('SPEECH_REGION'))

# Define the phrase to be spoken
phrase = "Hello, I'm a chatbot. How can I help you today? You can interrupt me whenever you want"

async def listen_for_user_input():
    speech_config.speech_recognition_language="en-US-ChristopherNeural"
    audio_config = speechsdk.audio.AudioConfig(use_default_microphone=True)
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config)

    result = await speech_recognizer.start_continuous_recognition_async()
    if result.reason == speechsdk.ResultReason.RecognizedSpeech:
        print("Recognized: {}".format(result.text))
        await speech_recognizer.stop_continuous_recognition_async()
    pass

async def speak_phrase(phrase):
    audio_config = speechsdk.audio.AudioOutputConfig(use_default_speaker=True)
    synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config)
    # Speak the defined phrase
    result = await synthesizer.speak_text_async(phrase)
    if result.reason == speechsdk.ResultReason.Canceled:
        cancellation_details = result.cancellation_details
        print("Azure Speech synthesis canceled: {}".format(cancellation_details.reason))
       

# Start the program
async def main():
    task1 = asyncio.create_task(speak_phrase(phrase))
    task2 = asyncio.create_task(listen_for_user_input())
    done, pending = await asyncio.wait([task1, task2], return_when=asyncio.FIRST_COMPLETED)
    for task in pending:
        task.cancel()

解决方案

问题分析

  1. 异步适配错误:Azure Speech SDK的async方法返回的是Future对象,不是原生asyncio可直接await的协程,需要用asyncio.wrap_future转换。
  2. 连续识别逻辑错误:start_continuous_recognition_async只是启动识别,不会直接返回识别结果,需要通过注册回调函数获取。
  3. 合成中断逻辑缺失:没有在检测到用户语音时主动停止正在进行的语音合成。

修正后的代码

import asyncio
import os
import azure.cognitiveservices.speech as speechsdk

# 配置Azure Speech服务
speech_config = speechsdk.SpeechConfig(subscription=os.environ.get('SPEECH_KEY'), region=os.environ.get('SPEECH_REGION'))
phrase = "Hello, I'm a chatbot. How can I help you today? You can interrupt me whenever you want"

# 全局变量用于控制合成中断
interrupt_triggered = False

def handle_recognized_result(evt):
    global interrupt_triggered
    if evt.result.reason == speechsdk.ResultReason.RecognizedSpeech:
        print(f"Recognized: {evt.result.text}")
        interrupt_triggered = True

async def listen_for_user_input():
    audio_config = speechsdk.audio.AudioConfig(use_default_microphone=True)
    speech_recognizer = speechsdk.SpeechRecognizer(
        speech_config=speech_config,
        audio_config=audio_config,
        speech_recognition_language="en-US"
    )
    # 注册识别结果回调
    speech_recognizer.recognized.connect(handle_recognized_result)
    
    # 启动连续识别
    await asyncio.wrap_future(speech_recognizer.start_continuous_recognition_async())
    
    # 等待中断触发
    while not interrupt_triggered:
        await asyncio.sleep(0.1)
    
    # 停止识别
    await asyncio.wrap_future(speech_recognizer.stop_continuous_recognition_async())

async def speak_phrase(phrase):
    global interrupt_triggered
    audio_config = speechsdk.audio.AudioOutputConfig(use_default_speaker=True)
    synthesizer = speechsdk.SpeechSynthesizer(
        speech_config=speech_config,
        audio_config=audio_config
    )
    
    # 启动合成任务
    synthesis_future = synthesizer.speak_text_async(phrase)
    
    # 轮询检查是否需要中断
    while not synthesis_future.done() and not interrupt_triggered:
        await asyncio.sleep(0.1)
    
    # 如果触发中断,停止合成
    if interrupt_triggered:
        synthesizer.stop_speaking_async()
        print("Speech synthesis interrupted by user input")
    
    result = await asyncio.wrap_future(synthesis_future)
    if result.reason == speechsdk.ResultReason.Canceled:
        cancellation_details = result.cancellation_details
        print(f"Azure Speech synthesis canceled: {cancellation_details.reason}")

async def main():
    global interrupt_triggered
    interrupt_triggered = False
    
    task1 = asyncio.create_task(speak_phrase(phrase))
    task2 = asyncio.create_task(listen_for_user_input())
    
    done, pending = await asyncio.wait([task1, task2], return_when=asyncio.FIRST_COMPLETED)
    for task in pending:
        task.cancel()
        try:
            await task
        except asyncio.CancelledError:
            pass

if __name__ == "__main__":
    asyncio.run(main())

关键改进点

  • 使用asyncio.wrap_future将SDK的Future对象转换为asyncio兼容的协程。
  • 通过回调函数handle_recognized_result实时获取语音识别结果,触发中断标记。
  • 在语音合成任务中轮询中断标记,一旦触发就调用stop_speaking_async停止合成。
  • 修复了连续语音识别的逻辑,确保持续监听用户输入。

内容的提问来源于stack exchange,提问作者Fabian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 03:55:16