You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于LiveKit的多模态Agent房间内重复发声问题求助

问题描述

我用LiveKit结合OpenAI RealtimeModel构建了多模态Agent,用于虚拟房间内的交互式面试,核心需求如下:

  • Agent通过音频实时对话(modalities=["audio"])
  • 到指定时间自动断开Agent
  • 断开前发送1分钟预警消息

✅ 正常运行功能

  • Agent可加入房间并响应语音
  • 转录记录(transcript)准确,用户与Agent消息均正确记录
  • 面试可按计划结束

❌ 问题现象

通过LiveKit Playground进入房间时,会听到Agent的语音两次播放,且两次响应内容不同,但转录记录仅显示每条消息的一个实例。


相关代码

面试逻辑代码

async def end_interview_at_time(ctx: JobContext, session, end_time: datetime, warning_time: datetime, my_shutdown_hook):
    """Task that handles interview timing and ending"""
    warning_sent = False
    
    while True:
        current_time = datetime.now()
        
        # Check if we need to send warning
        if not warning_sent and current_time >= warning_time:
            logger.info("Warning time reached, informing LLM to wrap up")
            wrap_up_instruction = "You have 1 minute remaining. Please wrap up the by asking any final questions conclude."
            await send_message(ctx, wrap_up_instruction)
            warning_sent = True
            logger.info("Wrap-up instruction sent to LLM")
        
        # Check if interview should end
        if current_time >= end_time:
            logger.info("duration completed, shutting down")
            # Add shutdown hook before ending the session
            ctx.add_shutdown_callback(my_shutdown_hook)
            # Clean up the room
            print(f"Cleaning up the room {ctx.room.name} \n\n")
            await room_manager.cleanup_room(ctx.room.name)
            # Shutdown the agent
            ctx.shutdown(reason="Interview duration completed")
            break
            
        await asyncio.sleep(10)

async def entrypoint(ctx: JobContext):
    metadata = json.loads(ctx.job.metadata)
    logger.info(f"\n\n job metadata type: {type(metadata)}")
    logger.info(f"\n\n job metadata: {metadata}")
    
    async def my_shutdown_hook():
        print(json.dumps(transcript, indent=2))
        
        # Store conversation in database           
    
    try:
        
        await ctx.connect(auto_subscribe=AutoSubscribe.SUBSCRIBE_ALL)
        await ctx.wait_for_participant()
        
        # Set up timer
        start_time = datetime.now()
        end_time = start_time + timedelta(minutes=metadata.get("duration_minutes") + 0.25)
        warning_time = end_time - timedelta(minutes=1)
        
        model = openai.realtime.RealtimeModel(
            instructions=prompts.get_system_prompt(),
            voice="shimmer",
            temperature=0.8,
            modalities=["audio", "text"]
        )
        agent = MultimodalAgent(
            model=model,
        )
        agent.start(ctx.room)
        
        session = model.sessions[0]
        
        # Start with welcome message
        welcome_message = VisaInterviewPrompts.get_welcome_message(metadata.get("visa_type"))
        # await send_voice_message(session, welcome_message)
        await send_message(ctx, welcome_message)
        
        
        @agent.on("user_speech_committed")
        def on_user_speech_committed(msg: llm.ChatMessage):
            # print(msg)
            # if (msg.content is not None and isinstance(msg.content, list)):
            #     msg = "\n".join("[image]" if isinstance(x, llm.ChatImage) else x for x in msg)
            transcript.append({
                "timestamp": get_formatted_timestamp(interview_start_time),
                "role": "applicant",
                "content": msg
            })
            handle_query(msg)
        
        @agent.on("agent_speech_committed")
        def on_agent_speech_committed(msg: llm.ChatMessage):
            # print("INTERVIEWER: ", msg.content)
            transcript.append({
                "timestamp": get_formatted_timestamp(interview_start_time),
                "role": "interviewer",
                "content": msg
            })
            
        def handle_query(msg: llm.ChatMessage):
            session.conversation.item.create(
                llm.ChatMessage(
                    role="user",
                    content=msg
                )
            )
            session.response.create()
        
        # Start the interview timer task
        timer_task = asyncio.create_task(end_interview_at_time(ctx, session, end_time, warning_time, my_shutdown_hook))
        
        # Wait for the timer task to complete (which will happen when the interview ends)
        await timer_task
        
    except Exception as e:
        logger.error(f"Error in main interview loop: {e}")
        try:
            ctx.add_shutdown_callback(my_shutdown_hook)
            # Clean up the room
            await room_manager.cleanup_room(ctx.room.name)
            # Shutdown the agent
            ctx.shutdown(reason="Error Encountered")
        except:
            pass
    
if __name__ == "__main__":
    cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint, agent_name="visa-interview-agent"))

Token生成与Agent调度代码

# token creation and dispatching the agent
token = (
                AccessToken(
                    self.api_key,
                    self.api_secret
                )
                .with_identity(participant_name)
                .with_grants(VideoGrants(room_join=True, room=room_name))
                .with_room_config(
                    RoomConfiguration(
                        agents=[
                            RoomAgentDispatch(
                                agent_name=self.agent_name, 
                                metadata=json.dumps(token_metadata),
                            )
                        ]
                    )
                )
                .to_jwt()
            )

解决方案

问题根源

出现重复且不同的语音,核心原因是用户消息被重复提交到OpenAI Realtime,生成了两个不同的响应,同时存在模态配置冲突导致的双重输出路径:

  1. MultimodalAgent内置了用户语音到OpenAI的提交逻辑,代码中又手动通过handle_query提交了一次,导致同一条用户消息触发两次LLM响应(因temperature=0.8,每次生成内容不同)
  2. model的modalities配置为["audio", "text"],同时存在send_message手动发送音频的逻辑,导致双重播放通道

修复步骤

  1. 移除重复的消息提交逻辑
    删除handle_query函数,以及user_speech_committed回调中对它的调用,依赖MultimodalAgent的默认逻辑处理用户语音到OpenAI的提交。

  2. 修正模态配置
    将RealtimeModel的modalities改为["audio"],符合核心需求,避免文本/音频双输出通道冲突。

  3. 统一欢迎消息发送方式
    通过OpenAI Session发送欢迎消息,避免send_message手动发送音频导致的重复播放。

修改后的核心代码片段

async def entrypoint(ctx: JobContext):
    metadata = json.loads(ctx.job.metadata)
    logger.info(f"\n\n job metadata type: {type(metadata)}")
    logger.info(f"\n\n job metadata: {metadata}")
    
    async def my_shutdown_hook():
        print(json.dumps(transcript, indent=2))
        
        # Store conversation in database           
    
    try:
        
        await ctx.connect(auto_subscribe=AutoSubscribe.SUBSCRIBE_ALL)
        await ctx.wait_for_participant()
        
        # Set up timer
        start_time = datetime.now()
        end_time = start_time + timedelta(minutes=metadata.get("duration_minutes") + 0.25)
        warning_time = end_time - timedelta(minutes=1)
        
        # 修正:仅使用audio模态
        model = openai.realtime.RealtimeModel(
            instructions=prompts.get_system_prompt(),
            voice="shimmer",
            temperature=0.8,
            modalities=["audio"]
        )
        agent = MultimodalAgent(
            model=model,
        )
        agent.start(ctx.room)
        
        session = model.sessions[0]
        
        # 修正:通过OpenAI Session发送欢迎消息,统一输出路径
        welcome_message = VisaInterviewPrompts.get_welcome_message(metadata.get("visa_type"))
        session.conversation.item.create(
            llm.ChatMessage(
                role="system",
                content=welcome_message
            )
        )
        session.response.create()
        
        # 仅保留转录逻辑,移除手动提交消息的代码
        @agent.on("user_speech_committed")
        def on_user_speech_committed(msg: llm.ChatMessage):
            transcript.append({
                "timestamp": get_formatted_timestamp(interview_start_time),
                "role": "applicant",
                "content": msg
            })
        
        @agent.on("agent_speech_committed")
        def on_agent_speech_committed(msg: llm.ChatMessage):
            transcript.append({
                "timestamp": get_formatted_timestamp(interview_start_time),
                "role": "interviewer",
                "content": msg
            })
            
        # Start the interview timer task
        timer_task = asyncio.create_task(end_interview_at_time(ctx, session, end_time, warning_time, my_shutdown_hook))
        
        # Wait for the timer task to complete (which will happen when the interview ends)
        await timer_task
        
    except Exception as e:
        logger.error(f"Error in main interview loop: {e}")
        try:
            ctx.add_shutdown_callback(my_shutdown_hook)
            # Clean up the room
            await room_manager.cleanup_room(ctx.room.name)
            # Shutdown the agent
            ctx.shutdown(reason="Error Encountered")
        except:
            pass

额外检查

确认send_message函数的实现,如果它是将文本直接转成语音发送到LiveKit房间,建议完全移除该函数的使用,统一通过OpenAI Realtime生成并输出语音,避免双重播放路径。


内容的提问来源于stack exchange,提问作者Fatima Sayeed Amani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 12:04:55