You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gemini Live音频转Vonage WebSocket随机倍速问题排查求助

问题分析与修复方案

可能原因分析

1. 跨线程异步调度冲突

当前实现采用threading.Thread运行asyncio事件循环,同时用线程安全的Queue传递音频数据。在低资源机器上,线程切换延迟会导致:

  • audio_sender_task中的await asyncio.sleep(0.02)无法保证精确的20ms间隔,实际延迟可能更长
  • 当延迟恢复后,缓冲区积累的大量音频会被一次性发送,导致Vonage端接收速率远超正常播放速度,表现为音频加速

2. 缓冲区缺乏流量控制

现有的AudioOutputBuffer仅实现了初始触发阈值,没有基于播放速率的流量限制:

  • 当Gemini一次性返回大段音频时,缓冲区瞬间堆满
  • audio_sender_task会一次性取出所有可用chunk发送,破坏了音频的时间连续性

3. 音频重采样性能瓶颈

使用pydub(依赖ffmpeg)进行重采样,在低资源机器上该操作CPU开销大:

  • 重采样延迟导致音频数据积压在缓冲区
  • 积压的数据后续集中发送时出现加速现象

4. WebSocket发送同步机制不足

虽然使用了asyncio.Lock保护WebSocket发送,但flask_sock的ws.send本质是同步操作,高负载下会导致:

  • audio_sender_task被阻塞,无法按计划发送音频
  • 阻塞解除后批量发送积压的音频chunk

潜在修复方案

1. 替换线程间队列为异步队列

将Queue改为asyncio.Queue,避免跨线程调度的开销,直接在asyncio事件循环内处理数据传递:

# 替换全局队列定义
audio_queue = asyncio.Queue()
text_queue = asyncio.Queue()

# 修改forward_vonage_to_gemini函数,直接异步取数据
async def forward_vonage_to_gemini():
    try:
        while True:
            data = await audio_queue.get()
            if isinstance(data, bytes):
                resampled_audio = resample_audio(data, input_rate=8000, target_rate=16000)
                await session.send(input={"data": resampled_audio, "mime_type": "audio/pcm"})
            elif data == "DONE":
                break
    except Exception as e:
        logger.error(f"Error forwarding Vonage to Gemini: {e}")

# 在handle_socket中用线程安全的方式向异步队列添加数据
loop.call_soon_threadsafe(audio_queue.put_nowait, data)

2. 基于播放时间的精确流量控制

重构audio_sender_task,严格按照音频chunk的播放时间间隔发送,确保发送速率与播放速率完全匹配:

async def audio_sender_task():
    """Send audio chunks at exact playback rate"""
    chunk_duration = 0.02  # 320字节对应8kHz 16位单声道的20ms音频
    while True:
        if output_buffer.started and len(output_buffer.buffer) >= 320:
            # 每次只取一个chunk,保证发送间隔精确
            chunk = bytes(output_buffer.buffer[:320])
            output_buffer.buffer = output_buffer.buffer[320:]
            output_buffer.total_sent += 320
            
            async with ws_lock:
                vonage_ws.send(chunk)
                stats['chunks_sent'] += 1
            
            # 等待对应播放时长,确保速率匹配
            await asyncio.sleep(chunk_duration)
        else:
            # 缓冲区不足时短等待,避免空转占用CPU
            await asyncio.sleep(0.001)

3. 优化音频重采样性能

改用numpy实现轻量级重采样,减少CPU占用:

import numpy as np

def resample_audio(data: bytes, input_rate=8000, target_rate=16000):
    """Fast resampling using numpy linear interpolation"""
    # 字节转numpy数组(16位PCM)
    audio_np = np.frombuffer(data, dtype=np.int16)
    # 计算重采样比例
    resample_ratio = target_rate / input_rate
    # 生成新的采样点
    new_length = int(len(audio_np) * resample_ratio)
    time_old = np.linspace(0, len(audio_np)-1, len(audio_np))
    time_new = np.linspace(0, len(audio_np)-1, new_length)
    # 线性插值重采样
    resampled_np = np.interp(time_new, time_old, audio_np).astype(np.int16)
    return resampled_np.tobytes()

def resample_audio_down(data: bytes, input_rate=24000, target_rate=8000):
    return resample_audio(data, input_rate, target_rate)

4. 添加缓冲区溢出保护

修改AudioOutputBuffer,设置最大缓冲区大小,避免数据积压过多:

class AudioOutputBuffer:
    """Buffer with flow control to prevent overflow"""
    def __init__(self, target_buffer_ms=100, max_buffer_ms=500, sample_rate=8000):
        self.buffer = bytearray()
        self.target_buffer_size = int((target_buffer_ms / 1000) * sample_rate * 2)
        # 限制最大缓冲区为500ms音频,避免积压过多
        self.max_buffer_size = int((max_buffer_ms / 1000) * sample_rate * 2)
        self.sample_rate = sample_rate
        self.started = False
        self.total_received = 0
        self.total_sent = 0
        
    def add_audio(self, audio_data):
        new_buffer_len = len(self.buffer) + len(audio_data)
        if new_buffer_len > self.max_buffer_size:
            # 丢弃旧数据,保留最新的音频
            discard_len = new_buffer_len - self.max_buffer_size
            self.buffer = self.buffer[discard_len:]
            logger.warning(f"Buffer overflow, discarded {discard_len} bytes")
        self.buffer.extend(audio_data)
        self.total_received += len(audio_data)
        
        if not self.started and len(self.buffer) >= self.target_buffer_size:
            self.started = True
            logger.info(f"Buffer ready, starting output. Size: {len(self.buffer)} bytes")

5. 优化asyncio事件循环配置

在低资源机器上,使用uvloop替代默认事件循环,提升异步性能:

# 安装uvloop:pip install uvloop
import uvloop

if __name__ == "__main__":
    print("Starting Gemini-Vonage Bridge Server")
    # 使用uvloop提升异步性能
    uvloop.install()
    app.run(port=8000, debug=False)  # 生产环境关闭debug模式

内容的提问来源于stack exchange,提问作者Jonathon Roscoe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 20:44:53