Gemini Live音频转Vonage WebSocket随机倍速问题排查求助
问题分析与修复方案
可能原因分析
1. 跨线程异步调度冲突
当前实现采用threading.Thread运行asyncio事件循环,同时用线程安全的Queue传递音频数据。在低资源机器上,线程切换延迟会导致:
audio_sender_task中的await asyncio.sleep(0.02)无法保证精确的20ms间隔,实际延迟可能更长- 当延迟恢复后,缓冲区积累的大量音频会被一次性发送,导致Vonage端接收速率远超正常播放速度,表现为音频加速
2. 缓冲区缺乏流量控制
现有的AudioOutputBuffer仅实现了初始触发阈值,没有基于播放速率的流量限制:
- 当Gemini一次性返回大段音频时,缓冲区瞬间堆满
audio_sender_task会一次性取出所有可用chunk发送,破坏了音频的时间连续性
3. 音频重采样性能瓶颈
使用pydub(依赖ffmpeg)进行重采样,在低资源机器上该操作CPU开销大:
- 重采样延迟导致音频数据积压在缓冲区
- 积压的数据后续集中发送时出现加速现象
4. WebSocket发送同步机制不足
虽然使用了asyncio.Lock保护WebSocket发送,但flask_sock的ws.send本质是同步操作,高负载下会导致:
audio_sender_task被阻塞,无法按计划发送音频- 阻塞解除后批量发送积压的音频chunk
潜在修复方案
1. 替换线程间队列为异步队列
将Queue改为asyncio.Queue,避免跨线程调度的开销,直接在asyncio事件循环内处理数据传递:
# 替换全局队列定义 audio_queue = asyncio.Queue() text_queue = asyncio.Queue() # 修改forward_vonage_to_gemini函数,直接异步取数据 async def forward_vonage_to_gemini(): try: while True: data = await audio_queue.get() if isinstance(data, bytes): resampled_audio = resample_audio(data, input_rate=8000, target_rate=16000) await session.send(input={"data": resampled_audio, "mime_type": "audio/pcm"}) elif data == "DONE": break except Exception as e: logger.error(f"Error forwarding Vonage to Gemini: {e}") # 在handle_socket中用线程安全的方式向异步队列添加数据 loop.call_soon_threadsafe(audio_queue.put_nowait, data)
2. 基于播放时间的精确流量控制
重构audio_sender_task,严格按照音频chunk的播放时间间隔发送,确保发送速率与播放速率完全匹配:
async def audio_sender_task(): """Send audio chunks at exact playback rate""" chunk_duration = 0.02 # 320字节对应8kHz 16位单声道的20ms音频 while True: if output_buffer.started and len(output_buffer.buffer) >= 320: # 每次只取一个chunk,保证发送间隔精确 chunk = bytes(output_buffer.buffer[:320]) output_buffer.buffer = output_buffer.buffer[320:] output_buffer.total_sent += 320 async with ws_lock: vonage_ws.send(chunk) stats['chunks_sent'] += 1 # 等待对应播放时长,确保速率匹配 await asyncio.sleep(chunk_duration) else: # 缓冲区不足时短等待,避免空转占用CPU await asyncio.sleep(0.001)
3. 优化音频重采样性能
改用numpy实现轻量级重采样,减少CPU占用:
import numpy as np def resample_audio(data: bytes, input_rate=8000, target_rate=16000): """Fast resampling using numpy linear interpolation""" # 字节转numpy数组(16位PCM) audio_np = np.frombuffer(data, dtype=np.int16) # 计算重采样比例 resample_ratio = target_rate / input_rate # 生成新的采样点 new_length = int(len(audio_np) * resample_ratio) time_old = np.linspace(0, len(audio_np)-1, len(audio_np)) time_new = np.linspace(0, len(audio_np)-1, new_length) # 线性插值重采样 resampled_np = np.interp(time_new, time_old, audio_np).astype(np.int16) return resampled_np.tobytes() def resample_audio_down(data: bytes, input_rate=24000, target_rate=8000): return resample_audio(data, input_rate, target_rate)
4. 添加缓冲区溢出保护
修改AudioOutputBuffer,设置最大缓冲区大小,避免数据积压过多:
class AudioOutputBuffer: """Buffer with flow control to prevent overflow""" def __init__(self, target_buffer_ms=100, max_buffer_ms=500, sample_rate=8000): self.buffer = bytearray() self.target_buffer_size = int((target_buffer_ms / 1000) * sample_rate * 2) # 限制最大缓冲区为500ms音频,避免积压过多 self.max_buffer_size = int((max_buffer_ms / 1000) * sample_rate * 2) self.sample_rate = sample_rate self.started = False self.total_received = 0 self.total_sent = 0 def add_audio(self, audio_data): new_buffer_len = len(self.buffer) + len(audio_data) if new_buffer_len > self.max_buffer_size: # 丢弃旧数据,保留最新的音频 discard_len = new_buffer_len - self.max_buffer_size self.buffer = self.buffer[discard_len:] logger.warning(f"Buffer overflow, discarded {discard_len} bytes") self.buffer.extend(audio_data) self.total_received += len(audio_data) if not self.started and len(self.buffer) >= self.target_buffer_size: self.started = True logger.info(f"Buffer ready, starting output. Size: {len(self.buffer)} bytes")
5. 优化asyncio事件循环配置
在低资源机器上,使用uvloop替代默认事件循环,提升异步性能:
# 安装uvloop:pip install uvloop import uvloop if __name__ == "__main__": print("Starting Gemini-Vonage Bridge Server") # 使用uvloop提升异步性能 uvloop.install() app.run(port=8000, debug=False) # 生产环境关闭debug模式
内容的提问来源于stack exchange,提问作者Jonathon Roscoe
相关产品推荐
相关产品推荐

