如何实现语音活动检测并在客户打断时停止Twilio语音机器人
实现Twilio媒体流的语音活动检测(VAD)打断机器人播报
要实现用户打断机器人播报的功能,核心是通过实时检测用户侧的语音活动,并在检测到语音时立即停止机器人的音频输出。结合你现有的Django+Twilio媒体流架构,可按以下步骤修改:
1. 安装VAD依赖
使用轻量高效的WebRTC VAD库来检测语音活动:
pip install webrtcvad
2. 修改WebSocket消费者逻辑
在你的TwilioWS类中添加VAD检测、状态跟踪和可控的音频发送逻辑:
2.1 导入必要模块
import webrtcvad import queue from threading import Event
2.2 初始化VAD和状态变量
在__init__方法中补充VAD相关初始化:
def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) self.groq_ai = None self.call_id = None self.streamSid = None self.transcriber = None self.tts = None self.thread_pool = concurrent.futures.ThreadPoolExecutor(max_workers=2) self.loop = asyncio.get_event_loop() # VAD核心配置 self.vad = webrtcvad.Vad() self.vad.set_mode(3) # 模式3为最严格检测,减少误判 self.is_bot_speaking = False # 标记机器人是否正在播报 self.audio_queue = queue.Queue() # 存储待发送的音频片段 self.stop_audio_event = Event() # 用于中断音频发送线程 self.voice_detected_count = 0 # 连续语音检测计数 self.VOICE_THRESHOLD = 3 # 连续3个片段检测到语音才判定为有效打断
2.3 新增VAD检测方法
实时处理用户侧的音频流,检测语音活动:
def detect_voice_activity(self, audio_chunk: bytes) -> None: # Twilio媒体流格式为16kHz、16位单声道PCM,每个片段为20ms(640字节) chunk_duration_ms = (len(audio_chunk) * 1000) // (16000 * 2) if chunk_duration_ms not in [10, 20, 30]: return # 检测当前片段是否包含语音 has_voice = self.vad.is_speech(audio_chunk, 16000) if has_voice: self.voice_detected_count += 1 # 连续检测到语音且机器人正在播报时,触发停止逻辑 if self.voice_detected_count >= self.VOICE_THRESHOLD and self.is_bot_speaking: self.stop_bot_speech() self.voice_detected_count = 0 else: self.voice_detected_count = 0
2.4 新增停止播报方法
def stop_bot_speech(self) -> None: self.is_bot_speaking = False self.stop_audio_event.set() # 清空音频队列,避免残留片段继续发送 while not self.audio_queue.empty(): try: self.audio_queue.get_nowait() except queue.Empty: pass
2.5 修改音频发送逻辑
改为队列+线程模式,支持随时中断:
def start_sending_audio(self, audio_data: bytes) -> None: # 将TTS音频分割为Twilio要求的20ms片段(640字节) chunk_size = 640 for i in range(0, len(audio_data), chunk_size): chunk = audio_data[i:i+chunk_size] # 补全最后一个片段的长度 if len(chunk) < chunk_size: chunk += b'\x00' * (chunk_size - len(chunk)) self.audio_queue.put(chunk) # 启动音频发送线程(仅当未在播报时) if not self.is_bot_speaking: self.is_bot_speaking = True self.stop_audio_event.clear() self.thread_pool.submit(self.send_audio_chunks) def send_audio_chunks(self) -> None: while self.is_bot_speaking and not self.stop_audio_event.is_set(): try: chunk = self.audio_queue.get(timeout=0.1) encoded_chunk = base64.b64encode(chunk).decode('UTF-8') self.send(json.dumps({ 'streamSid': self.streamSid, 'event': 'media', 'media': { 'payload': encoded_chunk, }, })) self.audio_queue.task_done() except queue.Empty: # 队列空了,结束播报 self.is_bot_speaking = False break
2.6 更新receive方法
在处理媒体流时优先执行VAD检测:
def receive(self, text_data: str) -> None: if not text_data: self.transcriber.disconnect() return data: dict = json.loads(text_data) event: str = data.get('event') if event == "start": self.streamSid: str = data['start']['streamSid'] self.call_id: str = data['start']['callSid'] if event == "media": audio_payload = base64.b64decode(data["media"]["payload"]) # 先检测用户语音,再处理转录 self.detect_voice_activity(audio_payload) self.handle_transcription(audio_payload, self.call_id) if event == "stop": self.transcriber.disconnect() self.stop_audio_event.set()
2.7 替换原音频发送方法
把原来的handle_transcribed_audio替换为新的队列发送逻辑:
def handle_transcribed_audio(self, audio_data): self.start_sending_audio(audio_data)
关键注意事项
- 确保TTS生成的音频格式为16kHz、16位单声道PCM,与Twilio媒体流格式一致,否则VAD检测和音频播放会异常。
- VAD模式可根据场景调整:模式0最宽松(适合弱语音),模式3最严格(减少误判)。
- 连续检测阈值
VOICE_THRESHOLD可根据实际通话环境调整,避免单次噪音误触发打断。
内容的提问来源于stack exchange,提问作者PAWAN BISHT
相关产品推荐
相关产品推荐

