You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Twilio双向媒体流中实现Dial转接人工号码?

问题背景与需求

我开发了一款基于Twilio双向媒体流(Bidirectional Media Stream)的语音应用,通过OpenAI Realtime API为来电者提供问答服务(代码灵感来源于Twilio官方集成示例)。核心的双向媒体流处理代码如下:

@app.websocket("/media-stream")
async def handle_media_stream(websocket: WebSocket):
    """Handle WebSocket connections between Twilio and OpenAI."""
    print("Client connected")
    await websocket.accept()

    async with websockets.connect(
        "wss://api.openai.com/v1/realtime?model=gpt-4o-realtime-preview-2024-10-01",
        extra_headers={
            "Authorization": f"Bearer {CONFIG.api_key}",
            "OpenAI-Beta": "realtime=v1",
        },
    ) as openai_ws:
        await initialize_session(openai_ws)

        # Connection specific state
        stream_sid = None
        latest_media_timestamp = 0
        last_assistant_item = None
        mark_queue = []
        response_start_timestamp_twilio = None

        async def receive_from_twilio():
            """Receive audio data from Twilio and send it to the OpenAI Realtime API."""
            nonlocal stream_sid, latest_media_timestamp
            try:
                async for message in websocket.iter_text():
                    data = json.loads(message)
                    if data["event"] == "media" and openai_ws.open:
                        latest_media_timestamp = int(data["media"]["timestamp"])
                        audio_append = {
                            "type": "input_audio_buffer.append",
                            "audio": data["media"]["payload"],
                        }
                        await openai_ws.send(json.dumps(audio_append))
                    elif data["event"] == "start":
                        stream_sid = data["start"]["streamSid"]
                        print(f"Incoming stream has started {stream_sid}")
                        response_start_timestamp_twilio = None  # noqa: F841
                        latest_media_timestamp = 0
                        last_assistant_item = None  # noqa: F841
                    elif data["event"] == "mark":
                        if mark_queue:
                            mark_queue.pop(0)
            except WebSocketDisconnect:
                print("Client disconnected.")
                if openai_ws.open:
                    await openai_ws.close()

        async def send_to_twilio():
            """Receive events from the OpenAI Realtime API, send audio back to Twilio."""
            nonlocal stream_sid, last_assistant_item, response_start_timestamp_twilio
            try:
                async for openai_message in openai_ws:
                    response = json.loads(openai_message)
                    response_type = response.get("type")
                    if response_type in CONFIG.log_event_types:
                        # print(f"Received event: {response['type']}", response)
                        logging.info(f"Received event: {response['type']}")

                    match response_type:
                        case "response.audio.delta":
                            if "delta" not in response:
                                continue

                            audio_payload = base64.b64encode(
                                base64.b64decode(response["delta"])
                            ).decode("utf-8")
                            audio_delta = {
                                "event": "media",
                                "streamSid": stream_sid,
                                "media": {"payload": audio_payload},
                            }
                            await websocket.send_json(audio_delta)

                            if response_start_timestamp_twilio is None:
                                response_start_timestamp_twilio = latest_media_timestamp
                                if CONFIG.show_timing_math:
                                    print(
                                        f"Setting start timestamp for new response: {response_start_timestamp_twilio}ms"
                                    )

                            # Update last_assistant_item safely
                            if response.get("item_id"):
                                last_assistant_item = response["item_id"]

                            await send_mark(websocket, stream_sid)

                        # Trigger an interruption. Your use case might work better using `input_audio_buffer.speech_stopped`, or combining the two.
                        case "input_audio_buffer.speech_started":
                            print("Speech started detected.")
                            if last_assistant_item:
                                print(
                                    f"Interrupting response with id: {last_assistant_item}"
                                )
                                await handle_speech_started_event()

                        case "response.function_call_arguments.done":
                            # https://platform.openai.com/docs/api-reference/realtime-server-events/response/function_call_arguments/done
                            # TODO: eventually migrate domain model to voice/
                            event = FunctionCallArgumentsEvent(**response)
                            logging.info(
                                f"Calling {event.name=} with {event.arguments=}"
                            )
                            await call_tool(
                                event.call_id,
                                event.name,
                                json.loads(event.arguments),
                                openai_ws,
                            )

            except Exception as e:
                traceback.print_exc()
                print(f"Error in send_to_twilio: {e}")

当前需求是:当来电者情绪不满时,将其转接至人工坐席号码,类似以下TwiML代码实现的Dial功能:

from twilio.twiml.voice_response import Dial, VoiceResponse, Say

response = VoiceResponse()
response.dial("111-111-1111") # dial out to human

但TwiML的Dial功能仅能在来电Webhook中使用,无法在双向媒体流的Webhook中调用,需要可行的变通方案实现媒体流中的程序化转接。

可行的变通方案

方案1:调用Twilio REST API更新呼叫路由

当检测到来电者需要转接时,直接通过Twilio的REST API修改当前呼叫的目标端点,触发Dial逻辑:

  1. 从Twilio发送的start事件中提取当前呼叫的Call SID(data["start"]["callSid"]),并在媒体流处理中保存该值
  2. 触发转接时,调用Twilio Calls API的更新接口,将呼叫的Url设置为返回Dial TwiML的后端端点
  3. 示例代码:
from twilio.rest import Client

# 初始化Twilio客户端
twilio_client = Client(TWILIO_ACCOUNT_SID, TWILIO_AUTH_TOKEN)

def transfer_to_agent(call_sid):
    # 指向返回Dial TwiML的后端端点
    twiml_url = "https://your-server.com/dial-agent"
    # 更新呼叫路由
    call = twilio_client.calls(call_sid).update(url=twiml_url, method="POST")
    print(f"Call transferred: {call.sid}")

对应的/dial-agent端点实现:

from twilio.twiml.voice_response import Dial, VoiceResponse
from fastapi import Response

@app.post("/dial-agent")
async def dial_agent():
    response = VoiceResponse()
    response.dial("111-111-1111")
    return Response(content=str(response), media_type="application/xml")

方案2:通过Call Control API发送重定向指令

利用Twilio的Call Control功能,直接发送redirect指令将呼叫导向Dial TwiML端点:

  1. 保存当前呼叫的Call SID
  2. 触发转接时,调用API将呼叫重定向到Dial端点,逻辑与方案1类似,本质是通过API强制修改呼叫的执行流程

方案3:终止媒体流并触发呼叫回退逻辑

  1. 触发转接时,先调用Twilio API更新呼叫的Url为Dial端点
  2. 主动关闭与Twilio的媒体流WebSocket连接,此时Twilio会立即执行新设置的Url中的TwiML逻辑,完成转接

关键注意事项

  • 必须确保Twilio API凭据拥有修改呼叫状态的权限
  • 转接前要妥善关闭媒体流连接,避免音频异常或呼叫状态混乱
  • 可结合OpenAI Realtime API的语音转文本结果,通过情绪分析自动触发转接逻辑

内容的提问来源于stack exchange,提问作者Noah Stebbins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 21:35:55