如何处理Azure Text-To-Speech输出以适配PyAudio音频流?
问题:Azure文本转语音输出通过PyAudio流式播放报错处理
我基于微软的文本转语音示例代码,尝试将Azure文本转语音(Text-To-Speech)实例的输出通过PyAudio流式传输到扬声器,但在Azure的回调函数def write中调用my_stream.write(audio_buffer)时出现如下错误:
my_stream.write(audio_buffer) File "/opt/homebrew/lib/python3.10/site-packages/pyaudio.py", line 589, in write pa.write_stream(self._stream, frames, num_frames, TypeError: argument 2 must be read-only bytes-like object, not memoryview
解决方法
- PyAudio的
write()方法仅接受只读bytes类型对象,而Azure返回的音频数据是memoryview类型,只需将其转换为bytes即可解决类型不匹配问题:将my_stream.write(audio_buffer)修改为my_stream.write(bytes(audio_buffer))。 - 另外修正原代码中的一处变量错误:合成完成提示中的
text变量未定义,替换为全局的my_text。 - 补充PyAudio资源释放逻辑,避免程序结束后残留未关闭的音频流。
完整修正代码
import azure.cognitiveservices.speech as speechsdk import os, sys, pyaudio pa = pyaudio.PyAudio() my_text = "My emotional experiences are varied, but mostly involve trying to find a balance between understanding others’ feelings and managing my own. I also explore the intersection of emotion and technology through affective computing and related research." voc_data = { 'channels': 1 if sys.platform == 'darwin' else 2, 'rate': 44100, 'width': pa.get_sample_size(pyaudio.paInt16), 'format': pyaudio.paInt16, 'frames': [] } my_stream = pa.open(format=voc_data['format'], channels=voc_data['channels'], rate=voc_data['rate'], output=True) speech_key = os.getenv('SPEECH_KEY') service_region = os.getenv('SPEECH_REGION') def speech_synthesis_to_push_audio_output_stream(): """performs speech synthesis and push audio output to a stream""" class PushAudioOutputStreamSampleCallback(speechsdk.audio.PushAudioOutputStreamCallback): """ Example class that implements the PushAudioOutputStreamCallback, which is used to show how to push output audio to a stream """ def __init__(self) -> None: super().__init__() self._audio_data = bytes(0) self._closed = False def write(self, audio_buffer: memoryview) -> int: """ The callback function which is invoked when the synthesizer has an output audio chunk to write out """ self._audio_data += audio_buffer # 转换memoryview为bytes后写入PyAudio流 my_stream.write(bytes(audio_buffer)) print("{} bytes received.".format(audio_buffer.nbytes)) return audio_buffer.nbytes def close(self) -> None: """ The callback function which is invoked when the synthesizer is about to close the stream. """ self._closed = True print("Push audio output stream closed.") # 释放PyAudio资源 my_stream.stop_stream() my_stream.close() pa.terminate() def get_audio_data(self) -> bytes: return self._audio_data def get_audio_size(self) -> int: return len(self._audio_data) # 创建语音配置实例 speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region) # 初始化自定义回调实例 stream_callback = PushAudioOutputStreamSampleCallback() # 创建推送音频输出流 push_stream = speechsdk.audio.PushAudioOutputStream(stream_callback) # 创建使用推送流的语音合成器 stream_config = speechsdk.audio.AudioOutputConfig(stream=push_stream) speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=stream_config) # 执行文本转语音 result = speech_synthesizer.speak_text_async(my_text).get() # 检查合成结果 if result.reason == speechsdk.ResultReason.SynthesizingAudioCompleted: # 修正未定义变量问题 print("Speech synthesized for text [{}], and the audio was written to output stream.".format(my_text)) elif result.reason == speechsdk.ResultReason.Canceled: cancellation_details = result.cancellation_details print("Speech synthesis canceled: {}".format(cancellation_details.reason)) if cancellation_details.reason == speechsdk.CancellationReason.Error: print("Error details: {}".format(cancellation_details.error_details)) # 释放结果资源 del result # 销毁合成器以关闭输出流 del speech_synthesizer print("Totally {} bytes received.".format(stream_callback.get_audio_size())) speech_synthesis_to_push_audio_output_stream()
内容的提问来源于stack exchange,提问作者bozolino
相关产品推荐
相关产品推荐

