You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python在Azure文本转语音中不播放语音保存流对象

解决方案:Azure Speech SDK 无播放合成音频并保存

直接使用synthesize_speech_to_stream_async方法即可实现无播放的音频合成,你遇到的格式错误是因为API更新后的类名变化,以下是修正后的完整实现:

核心代码示例

import azure.cognitiveservices.speech as speechsdk

# 初始化Speech服务配置
speech_config = speechsdk.SpeechConfig(subscription="你的订阅密钥", region="你的服务区域")

# 创建PCM音频格式(替代已弃用的AudioStreamFormat和不存在的PcmDataFormat)
# 使用默认PCM格式(16kHz采样率、16位深度、单声道)
pcm_format = speechsdk.audio.AudioStreamWaveFormat.get_default_pcm_format()

# 执行无播放的音频合成
synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config)
result_future = synthesizer.synthesize_speech_to_stream_async(
    text="这里替换成书籍的文本内容",
    audio_format=pcm_format
)

# 处理合成结果
result = result_future.get()
if result.reason == speechsdk.ResultReason.SynthesizingAudioCompleted:
    # 将音频数据写入WAV文件
    with open("book_audio.wav", "wb") as output_file:
        output_file.write(result.audio_data)
    print("音频文件保存成功")
else:
    print(f"合成失败:{result.error_details}")

关键说明

  • 弃用的AudioStreamFormat已被AudioStreamWaveFormat替代,无需再使用旧类。
  • 不存在PcmDataFormat属性,直接通过AudioStreamWaveFormat.get_default_pcm_format()获取标准PCM格式,或通过构造函数自定义参数(如采样率、声道数):
    # 自定义PCM格式示例:24kHz采样率、16位深度、双声道
    custom_pcm_format = speechsdk.audio.AudioStreamWaveFormat(
        samples_per_second=24000,
        bits_per_sample=16,
        channels=2
    )
    
  • synthesize_speech_to_stream_async方法不会触发本地语音播放,直接返回包含音频数据的合成结果,无需依赖播放流程获取流对象。

内容的提问来源于stack exchange,提问作者bobsmith76

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 01:34:57