Azure Speech麦克风连续语音识别:循环退出与终止方式咨询
Azure Speech Service 连续语音识别退出及交互方式问题解答
1. 无需手动输入命令停止连续语音识别
你的当前代码中,done_talking仅在会话停止或取消事件触发时才会设为True,但这两个事件不会主动触发,导致while循环无法退出。可以通过以下两种无手动输入的方式解决:
方式一:监听系统中断信号(如Ctrl+C)
添加信号处理逻辑,捕获终端中断信号时停止识别并退出循环:
import signal def cont_speech_to_text(): done_talking=False def stop_cb(evt): print('会话事件: {}'.format(evt)) nonlocal done_talking done_talking = True speech_recognizer.stop_continuous_recognition() # 新增信号处理函数 def handle_signal(signum, frame): nonlocal done_talking done_talking = True speech_recognizer.stop_continuous_recognition() print("\n已停止语音识别") signal.signal(signal.SIGINT, handle_signal) # 原有事件绑定逻辑 speech_recognizer.recognized.connect(recognised_speech) speech_recognizer.session_started.connect(lambda evt: print('SESSION STARTED: {}'.format(evt))) speech_recognizer.session_stopped.connect(lambda evt: print('SESSION STOPPED {}'.format(evt))) speech_recognizer.canceled.connect(lambda evt: print('CANCELED {}'.format(evt))) speech_recognizer.session_stopped.connect(stop_cb) speech_recognizer.canceled.connect(stop_cb) speech_recognizer.start_continuous_recognition() while not done_talking: time.sleep(.5)
方式二:检测静默时长自动停止
记录最后一次语音识别的时间,若超过设定的静默阈值(如10秒)则自动终止识别:
def cont_speech_to_text(): done_talking=False last_recognized_time = time.time() SILENCE_TIMEOUT = 10 # 静默10秒后停止 def recognised_speech(evt): nonlocal last_recognized_time last_recognized_time = time.time() print(f"You: {evt.result.text}") def stop_cb(evt): print('会话事件: {}'.format(evt)) nonlocal done_talking done_talking = True speech_recognizer.stop_continuous_recognition() # 原有事件绑定逻辑 speech_recognizer.recognized.connect(recognised_speech) speech_recognizer.session_started.connect(lambda evt: print('SESSION STARTED: {}'.format(evt))) speech_recognizer.session_stopped.connect(lambda evt: print('SESSION STOPPED {}'.format(evt))) speech_recognizer.canceled.connect(lambda evt: print('CANCELED {}'.format(evt))) speech_recognizer.session_stopped.connect(stop_cb) speech_recognizer.canceled.connect(stop_cb) speech_recognizer.start_continuous_recognition() while not done_talking: current_time = time.time() if current_time - last_recognized_time > SILENCE_TIMEOUT: print("检测到长时间静默,停止识别") done_talking = True speech_recognizer.stop_continuous_recognition() time.sleep(.5)
2. 通过语音命令/长停顿终止连续识别
方式一:关键词触发停止
在每次识别结果中检查是否包含指定停止关键词(如“退出”“停止识别”),匹配时主动终止:
def recognised_speech(evt): text = evt.result.text print(f"You: {text}") # 检查停止关键词 if any(keyword in text.lower() for keyword in ["退出", "停止识别", "结束"]): print("检测到停止命令,终止识别") speech_recognizer.stop_continuous_recognition()
方式二:长停顿触发停止
配置Azure Speech的端点检测参数,当检测到设定时长的静默时自动结束会话:
# 在创建speech_config后添加以下配置 # 设置结束静默超时(单位:毫秒),连续识别模式下静默超过此时间触发会话停止 speech_config.set_property(speechsdk.PropertyId.SpeechServiceConnection_EndSilenceTimeoutMs, "5000") # 5秒静默
配置后,麦克风静默超过5秒时会触发session_stopped事件,进而调用你的stop_cb函数停止循环。
3. 语音机器人的交互方式不止两种
你提到的单次识别(<15秒)和连续识别每次 utterance 后交互只是基础模式,Azure Speech Service还支持更多灵活的交互方式:
- 唤醒词触发识别:使用Keyword Recognition功能,仅当用户说出指定唤醒词(如“嘿,小助手”)时才激活识别,降低资源消耗。
- 会话式转录:针对多人对话场景,可实时区分说话人并转录完整对话,支持长时间持续会话。
- 实时双向语音交互:结合语音合成与识别,实现机器人实时响应,同时支持用户打断机器人语音输出的机制。
- 意图驱动的对话:配合Language Service的意图识别能力,自动理解用户需求并触发对应操作,无需手动处理每一句utterance。
内容的提问来源于stack exchange,提问作者erjvd
相关产品推荐
相关产品推荐

