You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu下Google Cloud实时语音转文本API自监听循环问题求助

解决Ubuntu聊天机器人麦克风拾音自触发无限循环问题

核心问题是机器人播放语音时,麦克风捕捉到自身输出导致重复识别,以下是两种可靠的代码层面解决方案,无需依赖系统静音命令:

方案一:直接控制麦克风音频流启停

通过暂停/恢复PyAudio的输入流,在机器人说话时切断麦克风输入,结束后恢复。

修改步骤:

  1. 给MicrophoneStream类添加暂停和恢复流的方法:
class MicrophoneStream(object):
    # ... 原有__init__、__enter__、__exit__等方法不变

    def pause_stream(self):
        """暂停麦克风输入流"""
        if self._audio_stream and self._audio_stream.is_active():
            self._audio_stream.stop_stream()

    def resume_stream(self):
        """恢复麦克风输入流"""
        if self._audio_stream and not self._audio_stream.is_active():
            self._audio_stream.start_stream()
  1. 修改listen_print_loop函数,在调用speak前后控制流状态(需传入MicrophoneStream实例):
def listen_print_loop(responses, stream):  # 新增stream参数
    global chat_log
    num_chars_printed = 0
    for response in responses:
        # ... 原有非final结果处理逻辑不变

        else:
            # ... 原有转录结果处理、退出判断逻辑不变

            question = transcript
            answer = ask(question, chat_log)
            chat_log = keepContext(question, answer, chat_log)
            
            # 说话前暂停麦克风
            stream.pause_stream()
            speak(answer)
            # 说话结束后恢复麦克风
            stream.resume_stream()
            
            print(answer)
            num_chars_printed = 0
  1. 调用listen_print_loop时传入MicrophoneStream实例:
with MicrophoneStream(RATE, CHUNK) as stream:
    responses = client.streaming_recognize(request, streaming_config=streaming_config)
    listen_print_loop(responses, stream)  # 传入stream实例

方案二:用线程事件过滤麦克风输入

通过线程事件标记机器人是否正在说话,在麦克风回调函数中跳过数据入队,避免无效输入进入识别流程。

修改步骤:

  1. 导入threading模块,给MicrophoneStream类添加事件控制:
import threading
import queue
import pyaudio

class MicrophoneStream(object):
    def __init__(self, rate, chunk):
        self._rate = rate
        self._chunk = chunk
        self._buff = queue.Queue()
        self.closed = True
        self._is_speaking = threading.Event()  # 标记是否正在播放语音

    # ... 原有__enter__、__exit__等方法不变

    def _fill_buffer(self, in_data, frame_count, time_info, status_flags):
        # 仅当机器人未说话时,才将麦克风数据加入缓冲区
        if not self._is_speaking.is_set():
            self._buff.put(in_data)
        return None, pyaudio.paContinue

    def start_speaking(self):
        """标记开始播放语音"""
        self._is_speaking.set()

    def stop_speaking(self):
        """标记结束播放语音"""
        self._is_speaking.clear()
  1. 修改listen_print_loop函数,在speak前后触发事件:
def listen_print_loop(responses, stream):  # 新增stream参数
    global chat_log
    num_chars_printed = 0
    for response in responses:
        # ... 原有逻辑不变

        else:
            # ... 原有转录结果处理、退出判断逻辑不变

            question = transcript
            answer = ask(question, chat_log)
            chat_log = keepContext(question, answer, chat_log)
            
            stream.start_speaking()
            speak(answer)
            stream.stop_speaking()
            
            print(answer)
            num_chars_printed = 0
  1. 调用listen_print_loop时传入MicrophoneStream实例,与方案一调用方式一致。

方案对比

  • 方案一更直接,彻底切断麦克风输入,适合对识别精度要求高的场景;
  • 方案二保留流运行,只是过滤数据,适合需要快速恢复输入的场景,避免流启停的开销。

内容的提问来源于stack exchange,提问作者Jack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 12:33:19