基于LiveKit的多模态Agent房间内重复发声问题求助
问题描述
我用LiveKit结合OpenAI RealtimeModel构建了多模态Agent,用于虚拟房间内的交互式面试,核心需求如下:
- Agent通过音频实时对话(
modalities=["audio"]) - 到指定时间自动断开Agent
- 断开前发送1分钟预警消息
✅ 正常运行功能
- Agent可加入房间并响应语音
- 转录记录(transcript)准确,用户与Agent消息均正确记录
- 面试可按计划结束
❌ 问题现象
通过LiveKit Playground进入房间时,会听到Agent的语音两次播放,且两次响应内容不同,但转录记录仅显示每条消息的一个实例。
相关代码
面试逻辑代码
async def end_interview_at_time(ctx: JobContext, session, end_time: datetime, warning_time: datetime, my_shutdown_hook): """Task that handles interview timing and ending""" warning_sent = False while True: current_time = datetime.now() # Check if we need to send warning if not warning_sent and current_time >= warning_time: logger.info("Warning time reached, informing LLM to wrap up") wrap_up_instruction = "You have 1 minute remaining. Please wrap up the by asking any final questions conclude." await send_message(ctx, wrap_up_instruction) warning_sent = True logger.info("Wrap-up instruction sent to LLM") # Check if interview should end if current_time >= end_time: logger.info("duration completed, shutting down") # Add shutdown hook before ending the session ctx.add_shutdown_callback(my_shutdown_hook) # Clean up the room print(f"Cleaning up the room {ctx.room.name} \n\n") await room_manager.cleanup_room(ctx.room.name) # Shutdown the agent ctx.shutdown(reason="Interview duration completed") break await asyncio.sleep(10) async def entrypoint(ctx: JobContext): metadata = json.loads(ctx.job.metadata) logger.info(f"\n\n job metadata type: {type(metadata)}") logger.info(f"\n\n job metadata: {metadata}") async def my_shutdown_hook(): print(json.dumps(transcript, indent=2)) # Store conversation in database try: await ctx.connect(auto_subscribe=AutoSubscribe.SUBSCRIBE_ALL) await ctx.wait_for_participant() # Set up timer start_time = datetime.now() end_time = start_time + timedelta(minutes=metadata.get("duration_minutes") + 0.25) warning_time = end_time - timedelta(minutes=1) model = openai.realtime.RealtimeModel( instructions=prompts.get_system_prompt(), voice="shimmer", temperature=0.8, modalities=["audio", "text"] ) agent = MultimodalAgent( model=model, ) agent.start(ctx.room) session = model.sessions[0] # Start with welcome message welcome_message = VisaInterviewPrompts.get_welcome_message(metadata.get("visa_type")) # await send_voice_message(session, welcome_message) await send_message(ctx, welcome_message) @agent.on("user_speech_committed") def on_user_speech_committed(msg: llm.ChatMessage): # print(msg) # if (msg.content is not None and isinstance(msg.content, list)): # msg = "\n".join("[image]" if isinstance(x, llm.ChatImage) else x for x in msg) transcript.append({ "timestamp": get_formatted_timestamp(interview_start_time), "role": "applicant", "content": msg }) handle_query(msg) @agent.on("agent_speech_committed") def on_agent_speech_committed(msg: llm.ChatMessage): # print("INTERVIEWER: ", msg.content) transcript.append({ "timestamp": get_formatted_timestamp(interview_start_time), "role": "interviewer", "content": msg }) def handle_query(msg: llm.ChatMessage): session.conversation.item.create( llm.ChatMessage( role="user", content=msg ) ) session.response.create() # Start the interview timer task timer_task = asyncio.create_task(end_interview_at_time(ctx, session, end_time, warning_time, my_shutdown_hook)) # Wait for the timer task to complete (which will happen when the interview ends) await timer_task except Exception as e: logger.error(f"Error in main interview loop: {e}") try: ctx.add_shutdown_callback(my_shutdown_hook) # Clean up the room await room_manager.cleanup_room(ctx.room.name) # Shutdown the agent ctx.shutdown(reason="Error Encountered") except: pass if __name__ == "__main__": cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint, agent_name="visa-interview-agent"))
Token生成与Agent调度代码
# token creation and dispatching the agent token = ( AccessToken( self.api_key, self.api_secret ) .with_identity(participant_name) .with_grants(VideoGrants(room_join=True, room=room_name)) .with_room_config( RoomConfiguration( agents=[ RoomAgentDispatch( agent_name=self.agent_name, metadata=json.dumps(token_metadata), ) ] ) ) .to_jwt() )
解决方案
问题根源
出现重复且不同的语音,核心原因是用户消息被重复提交到OpenAI Realtime,生成了两个不同的响应,同时存在模态配置冲突导致的双重输出路径:
MultimodalAgent内置了用户语音到OpenAI的提交逻辑,代码中又手动通过handle_query提交了一次,导致同一条用户消息触发两次LLM响应(因temperature=0.8,每次生成内容不同)model的modalities配置为["audio", "text"],同时存在send_message手动发送音频的逻辑,导致双重播放通道
修复步骤
移除重复的消息提交逻辑
删除handle_query函数,以及user_speech_committed回调中对它的调用,依赖MultimodalAgent的默认逻辑处理用户语音到OpenAI的提交。修正模态配置
将RealtimeModel的modalities改为["audio"],符合核心需求,避免文本/音频双输出通道冲突。统一欢迎消息发送方式
通过OpenAI Session发送欢迎消息,避免send_message手动发送音频导致的重复播放。
修改后的核心代码片段
async def entrypoint(ctx: JobContext): metadata = json.loads(ctx.job.metadata) logger.info(f"\n\n job metadata type: {type(metadata)}") logger.info(f"\n\n job metadata: {metadata}") async def my_shutdown_hook(): print(json.dumps(transcript, indent=2)) # Store conversation in database try: await ctx.connect(auto_subscribe=AutoSubscribe.SUBSCRIBE_ALL) await ctx.wait_for_participant() # Set up timer start_time = datetime.now() end_time = start_time + timedelta(minutes=metadata.get("duration_minutes") + 0.25) warning_time = end_time - timedelta(minutes=1) # 修正:仅使用audio模态 model = openai.realtime.RealtimeModel( instructions=prompts.get_system_prompt(), voice="shimmer", temperature=0.8, modalities=["audio"] ) agent = MultimodalAgent( model=model, ) agent.start(ctx.room) session = model.sessions[0] # 修正:通过OpenAI Session发送欢迎消息,统一输出路径 welcome_message = VisaInterviewPrompts.get_welcome_message(metadata.get("visa_type")) session.conversation.item.create( llm.ChatMessage( role="system", content=welcome_message ) ) session.response.create() # 仅保留转录逻辑,移除手动提交消息的代码 @agent.on("user_speech_committed") def on_user_speech_committed(msg: llm.ChatMessage): transcript.append({ "timestamp": get_formatted_timestamp(interview_start_time), "role": "applicant", "content": msg }) @agent.on("agent_speech_committed") def on_agent_speech_committed(msg: llm.ChatMessage): transcript.append({ "timestamp": get_formatted_timestamp(interview_start_time), "role": "interviewer", "content": msg }) # Start the interview timer task timer_task = asyncio.create_task(end_interview_at_time(ctx, session, end_time, warning_time, my_shutdown_hook)) # Wait for the timer task to complete (which will happen when the interview ends) await timer_task except Exception as e: logger.error(f"Error in main interview loop: {e}") try: ctx.add_shutdown_callback(my_shutdown_hook) # Clean up the room await room_manager.cleanup_room(ctx.room.name) # Shutdown the agent ctx.shutdown(reason="Error Encountered") except: pass
额外检查
确认send_message函数的实现,如果它是将文本直接转成语音发送到LiveKit房间,建议完全移除该函数的使用,统一通过OpenAI Realtime生成并输出语音,避免双重播放路径。
内容的提问来源于stack exchange,提问作者Fatima Sayeed Amani
相关产品推荐
相关产品推荐

