You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Twilio+FastAPI WebSocket发送Azure TTS音频报错:字符串索引必须为整数

问题:使用Twilio+Azure文本转语音向通话者发送音频时遇WebSocket发送错误

我正在用Twilio和Azure文本转语音(Text-to-Speech)实现向通话者发送可收听音频的功能,但调用ws.send函数时抛出错误:

Error sending message: string indices must be integers

我的send_raw_audio函数代码:

async def send_raw_audio(mulaw_bytes, ws:WebSocket, stream_sid):
    # Step 1: Encode the raw Mu-Law audio bytes to Base64
    print("Type of mulaw_bytes:", type(mulaw_bytes))
    print("Type of ws:", type(ws))
    print("Type of stream_sid:", type(stream_sid))

    base64_data = base64.b64encode(mulaw_bytes).decode("utf-8")

    message = {
        "event": "media",
        "streamSid": stream_sid,
        "media": {
            "payload": base64_data
        }
    }

    # Step 3: Send the message over WebSocket
    message_json = json.dumps(message)
    try:
        await ws.send(message_json)
        print("Message sent successfully:", message_json)
    except Exception as e:
        print("Error sending message:", e)
        raise

AzureSpeechSynthesizer类代码:

class AzureSpeechSynthesizer:
    key = os.environ['SPEECH_KEY']
    service_region = os.environ['SPEECH_REGION']

    def __init__(self, language: str, speech_synthesis_voice_name: str):
        self.speech_config = speechsdk.SpeechConfig(subscription=self.key, region=self.service_region, speech_recognition_language=language)
        self.speech_config.set_speech_synthesis_output_format(speechsdk.SpeechSynthesisOutputFormat.Raw8Khz8BitMonoMULaw)
        self.speech_config.speech_synthesis_voice_name = speech_synthesis_voice_name

    def text_to_raw_mulaw(self, text: str) -> bytes:
        synthesizer = speechsdk.SpeechSynthesizer(speech_config=self.speech_config, audio_config=None)
        result = synthesizer.speak_text_async(text).get()

        if result.reason != speechsdk.ResultReason.SynthesizingAudioCompleted:
            raise Exception(f"Synthesis failed: {result.reason}")

        return result.audio_data

音频数据状态:

  • 生成的base64编码音频数据格式正常
  • 将base64数据解码为线性PCM后,u-law音频波形也正常

更新

我发现不封装成send_raw_audio函数,直接在逻辑里发送音频就能成功返回给通话者,比如官方示例代码:

def echo(ws):
    app.logger.info("Connection accepted")
    # A lot of messages will be sent rapidly. We'll stop showing after the first one.
    has_seen_media = False
    message_count = 0
    while not ws.closed:
        message = ws.receive()
        if message is None:
            app.logger.info("No message received...")
            continue

        # Messages are a JSON encoded string
        data = json.loads(message)

        # Using the event type you can determine what type of message you are receiving
        if data['event'] == "connected":
            app.logger.info("Connected Message received: {}".format(message))
        if data['event'] == "start":
            app.logger.info("Start Message received: {}".format(message))
        if data['event'] == "media":
            if not has_seen_media:
                app.logger.info("Media message: {}".format(message))
                payload = data['media']['payload']
                app.logger.info("Payload is: {}".format(payload))
                chunk = base64.b64decode(payload)
                message = {
                   "event": "media",
                   "streamSid": stream_sid,
                   "media": {
                      "payload": base64_data
                    }
                }
                await ws.send_json(message) #this will work
                has_seen_media = True
        if data['event'] == "closed":
            app.logger.info("Closed Message received: {}".format(message))
            break
        message_count += 1

为什么直接发送可行,但封装成send_raw_audio函数就失败?

内容的提问来源于stack exchange,提问作者J7er

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 03:05:19