You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现语音活动检测并在客户打断时停止Twilio语音机器人

实现Twilio媒体流的语音活动检测(VAD)打断机器人播报

要实现用户打断机器人播报的功能,核心是通过实时检测用户侧的语音活动,并在检测到语音时立即停止机器人的音频输出。结合你现有的Django+Twilio媒体流架构,可按以下步骤修改:

1. 安装VAD依赖

使用轻量高效的WebRTC VAD库来检测语音活动:

pip install webrtcvad

2. 修改WebSocket消费者逻辑

在你的TwilioWS类中添加VAD检测、状态跟踪和可控的音频发送逻辑:

2.1 导入必要模块

import webrtcvad
import queue
from threading import Event

2.2 初始化VAD和状态变量

在__init__方法中补充VAD相关初始化:

def __init__(self, *args, **kwargs):
    super().__init__(*args, **kwargs)
    
    self.groq_ai = None
    self.call_id = None
    self.streamSid = None
    self.transcriber = None
    self.tts = None
    self.thread_pool = concurrent.futures.ThreadPoolExecutor(max_workers=2)
    self.loop = asyncio.get_event_loop()
    
    # VAD核心配置
    self.vad = webrtcvad.Vad()
    self.vad.set_mode(3)  # 模式3为最严格检测,减少误判
    self.is_bot_speaking = False  # 标记机器人是否正在播报
    self.audio_queue = queue.Queue()  # 存储待发送的音频片段
    self.stop_audio_event = Event()  # 用于中断音频发送线程
    self.voice_detected_count = 0  # 连续语音检测计数
    self.VOICE_THRESHOLD = 3  # 连续3个片段检测到语音才判定为有效打断

2.3 新增VAD检测方法

实时处理用户侧的音频流,检测语音活动:

def detect_voice_activity(self, audio_chunk: bytes) -> None:
    # Twilio媒体流格式为16kHz、16位单声道PCM,每个片段为20ms(640字节)
    chunk_duration_ms = (len(audio_chunk) * 1000) // (16000 * 2)
    if chunk_duration_ms not in [10, 20, 30]:
        return

    # 检测当前片段是否包含语音
    has_voice = self.vad.is_speech(audio_chunk, 16000)
    
    if has_voice:
        self.voice_detected_count += 1
        # 连续检测到语音且机器人正在播报时,触发停止逻辑
        if self.voice_detected_count >= self.VOICE_THRESHOLD and self.is_bot_speaking:
            self.stop_bot_speech()
            self.voice_detected_count = 0
    else:
        self.voice_detected_count = 0

2.4 新增停止播报方法

def stop_bot_speech(self) -> None:
    self.is_bot_speaking = False
    self.stop_audio_event.set()
    # 清空音频队列,避免残留片段继续发送
    while not self.audio_queue.empty():
        try:
            self.audio_queue.get_nowait()
        except queue.Empty:
            pass

2.5 修改音频发送逻辑

改为队列+线程模式,支持随时中断:

def start_sending_audio(self, audio_data: bytes) -> None:
    # 将TTS音频分割为Twilio要求的20ms片段(640字节)
    chunk_size = 640
    for i in range(0, len(audio_data), chunk_size):
        chunk = audio_data[i:i+chunk_size]
        # 补全最后一个片段的长度
        if len(chunk) < chunk_size:
            chunk += b'\x00' * (chunk_size - len(chunk))
        self.audio_queue.put(chunk)
    
    # 启动音频发送线程(仅当未在播报时)
    if not self.is_bot_speaking:
        self.is_bot_speaking = True
        self.stop_audio_event.clear()
        self.thread_pool.submit(self.send_audio_chunks)

def send_audio_chunks(self) -> None:
    while self.is_bot_speaking and not self.stop_audio_event.is_set():
        try:
            chunk = self.audio_queue.get(timeout=0.1)
            encoded_chunk = base64.b64encode(chunk).decode('UTF-8')
            self.send(json.dumps({
                'streamSid': self.streamSid,
                'event': 'media',
                'media': {
                    'payload': encoded_chunk,
                },
            }))
            self.audio_queue.task_done()
        except queue.Empty:
            # 队列空了,结束播报
            self.is_bot_speaking = False
            break

2.6 更新receive方法

在处理媒体流时优先执行VAD检测:

def receive(self, text_data: str) -> None:
    if not text_data:
        self.transcriber.disconnect()
        return

    data: dict = json.loads(text_data)
    event: str = data.get('event')

    if event == "start":
        self.streamSid: str = data['start']['streamSid']
        self.call_id: str = data['start']['callSid']

    if event == "media":
        audio_payload = base64.b64decode(data["media"]["payload"])
        # 先检测用户语音,再处理转录
        self.detect_voice_activity(audio_payload)
        self.handle_transcription(audio_payload, self.call_id)

    if event == "stop":
        self.transcriber.disconnect()
        self.stop_audio_event.set()

2.7 替换原音频发送方法

把原来的handle_transcribed_audio替换为新的队列发送逻辑:

def handle_transcribed_audio(self, audio_data):
    self.start_sending_audio(audio_data)

关键注意事项

  • 确保TTS生成的音频格式为16kHz、16位单声道PCM,与Twilio媒体流格式一致,否则VAD检测和音频播放会异常。
  • VAD模式可根据场景调整:模式0最宽松(适合弱语音),模式3最严格(减少误判)。
  • 连续检测阈值VOICE_THRESHOLD可根据实际通话环境调整,避免单次噪音误触发打断。

内容的提问来源于stack exchange,提问作者PAWAN BISHT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 06:49:55