You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何处理Azure Text-To-Speech输出以适配PyAudio音频流?

问题:Azure文本转语音输出通过PyAudio流式播放报错处理

我基于微软的文本转语音示例代码,尝试将Azure文本转语音(Text-To-Speech)实例的输出通过PyAudio流式传输到扬声器,但在Azure的回调函数def write中调用my_stream.write(audio_buffer)时出现如下错误:

my_stream.write(audio_buffer)
File "/opt/homebrew/lib/python3.10/site-packages/pyaudio.py", line 589, in write pa.write_stream(self._stream, frames, num_frames,
TypeError: argument 2 must be read-only bytes-like object, not memoryview

解决方法

  • PyAudio的write()方法仅接受只读bytes类型对象,而Azure返回的音频数据是memoryview类型,只需将其转换为bytes即可解决类型不匹配问题:将my_stream.write(audio_buffer)修改为my_stream.write(bytes(audio_buffer))。
  • 另外修正原代码中的一处变量错误:合成完成提示中的text变量未定义,替换为全局的my_text。
  • 补充PyAudio资源释放逻辑,避免程序结束后残留未关闭的音频流。

完整修正代码

import azure.cognitiveservices.speech as speechsdk
import os, sys, pyaudio
pa = pyaudio.PyAudio()

my_text = "My emotional experiences are varied, but mostly involve trying to find a balance between understanding others’ feelings and managing my own. I also explore the intersection of emotion and technology through affective computing and related research."

voc_data = {
    'channels': 1 if sys.platform == 'darwin' else 2,
    'rate': 44100,
    'width': pa.get_sample_size(pyaudio.paInt16),
    'format': pyaudio.paInt16,
    'frames': []
}

my_stream = pa.open(format=voc_data['format'],
                    channels=voc_data['channels'],
                    rate=voc_data['rate'],
                    output=True)

speech_key = os.getenv('SPEECH_KEY')
service_region = os.getenv('SPEECH_REGION')

def speech_synthesis_to_push_audio_output_stream():
    """performs speech synthesis and push audio output to a stream"""
    class PushAudioOutputStreamSampleCallback(speechsdk.audio.PushAudioOutputStreamCallback):
        """
        Example class that implements the PushAudioOutputStreamCallback, which is used to show
        how to push output audio to a stream
        """
        def __init__(self) -> None:
            super().__init__()
            self._audio_data = bytes(0)
            self._closed = False
        def write(self, audio_buffer: memoryview) -> int:
            """
            The callback function which is invoked when the synthesizer has an output audio chunk
            to write out
            """
            self._audio_data += audio_buffer
            # 转换memoryview为bytes后写入PyAudio流
            my_stream.write(bytes(audio_buffer))
            print("{} bytes received.".format(audio_buffer.nbytes))
            return audio_buffer.nbytes

        def close(self) -> None:
            """
            The callback function which is invoked when the synthesizer is about to close the
            stream.
            """
            self._closed = True
            print("Push audio output stream closed.")
            # 释放PyAudio资源
            my_stream.stop_stream()
            my_stream.close()
            pa.terminate()

        def get_audio_data(self) -> bytes:
            return self._audio_data

        def get_audio_size(self) -> int:
            return len(self._audio_data)

    # 创建语音配置实例
    speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)
    # 初始化自定义回调实例
    stream_callback = PushAudioOutputStreamSampleCallback()
    # 创建推送音频输出流
    push_stream = speechsdk.audio.PushAudioOutputStream(stream_callback)
    # 创建使用推送流的语音合成器
    stream_config = speechsdk.audio.AudioOutputConfig(stream=push_stream)
    speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=stream_config)

    # 执行文本转语音
    result = speech_synthesizer.speak_text_async(my_text).get()
    # 检查合成结果
    if result.reason == speechsdk.ResultReason.SynthesizingAudioCompleted:
        # 修正未定义变量问题
        print("Speech synthesized for text [{}], and the audio was written to output stream.".format(my_text))
    elif result.reason == speechsdk.ResultReason.Canceled:
        cancellation_details = result.cancellation_details
        print("Speech synthesis canceled: {}".format(cancellation_details.reason))
        if cancellation_details.reason == speechsdk.CancellationReason.Error:
            print("Error details: {}".format(cancellation_details.error_details))
    # 释放结果资源
    del result

    # 销毁合成器以关闭输出流
    del speech_synthesizer

    print("Totally {} bytes received.".format(stream_callback.get_audio_size()))

speech_synthesis_to_push_audio_output_stream()

内容的提问来源于stack exchange,提问作者bozolino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 00:52:40