如何用Python在Azure文本转语音中不播放语音保存流对象
解决方案:Azure Speech SDK 无播放合成音频并保存
直接使用synthesize_speech_to_stream_async方法即可实现无播放的音频合成,你遇到的格式错误是因为API更新后的类名变化,以下是修正后的完整实现:
核心代码示例
import azure.cognitiveservices.speech as speechsdk # 初始化Speech服务配置 speech_config = speechsdk.SpeechConfig(subscription="你的订阅密钥", region="你的服务区域") # 创建PCM音频格式(替代已弃用的AudioStreamFormat和不存在的PcmDataFormat) # 使用默认PCM格式(16kHz采样率、16位深度、单声道) pcm_format = speechsdk.audio.AudioStreamWaveFormat.get_default_pcm_format() # 执行无播放的音频合成 synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config) result_future = synthesizer.synthesize_speech_to_stream_async( text="这里替换成书籍的文本内容", audio_format=pcm_format ) # 处理合成结果 result = result_future.get() if result.reason == speechsdk.ResultReason.SynthesizingAudioCompleted: # 将音频数据写入WAV文件 with open("book_audio.wav", "wb") as output_file: output_file.write(result.audio_data) print("音频文件保存成功") else: print(f"合成失败:{result.error_details}")
关键说明
- 弃用的
AudioStreamFormat已被AudioStreamWaveFormat替代,无需再使用旧类。 - 不存在
PcmDataFormat属性,直接通过AudioStreamWaveFormat.get_default_pcm_format()获取标准PCM格式,或通过构造函数自定义参数(如采样率、声道数):# 自定义PCM格式示例:24kHz采样率、16位深度、双声道 custom_pcm_format = speechsdk.audio.AudioStreamWaveFormat( samples_per_second=24000, bits_per_sample=16, channels=2 ) synthesize_speech_to_stream_async方法不会触发本地语音播放,直接返回包含音频数据的合成结果,无需依赖播放流程获取流对象。
内容的提问来源于stack exchange,提问作者bobsmith76
相关产品推荐
相关产品推荐

