You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何禁用Azure TTS生成音频样本时的自动播放功能?

解决Azure Speech Service TTS禁用音频播放的问题

问题原因

直接初始化AudioOutputConfig()时,SDK没有明确的输出目标,会触发ValueError;而显式设置use_default_speaker=False时,SDK仍会默认调用系统音频输出设备,导致音频播放。

正确解决方案

使用SDK提供的NullAudioOutputStream,它会直接丢弃合成的音频数据,既不播放也无需写入文件。代码示例如下:

import os
import azure.cognitiveservices.speech as speechsdk

speech_config = speechsdk.SpeechConfig(
    subscription=os.environ.get('SPEECH_KEY'),
    region=os.environ.get('SPEECH_REGION')
)
# 创建空音频输出流,直接丢弃合成的音频数据
null_stream = speechsdk.audio.NullAudioOutputStream()
audio_config = speechsdk.audio.AudioOutputConfig(stream=null_stream)

# 初始化合成器并执行文本合成
speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config)
result = speech_synthesizer.speak_text_async("需要合成的目标文本").get()

# 可选:检查合成结果状态
if result.reason == speechsdk.ResultReason.SynthesizingAudioCompleted:
    print("文本合成完成")
elif result.reason == speechsdk.ResultReason.Canceled:
    cancellation_details = result.cancellation_details
    print(f"合成任务取消: {cancellation_details.reason}")
    if cancellation_details.reason == speechsdk.CancellationReason.Error:
        print(f"错误详情: {cancellation_details.error_details}")

补充:获取音频数据但不播放

如果需要在不播放的前提下获取合成的音频字节数据,可以自定义PullAudioOutputStream缓存数据到内存:

import os
import azure.cognitiveservices.speech as speechsdk

# 自定义拉取式音频流回调,用于缓存音频数据
class AudioBufferCallback(speechsdk.audio.PullAudioOutputStreamCallback):
    def __init__(self):
        super().__init__()
        self.audio_buffer = bytearray()

    def pull(self, buffer: memoryview) -> int:
        # 将音频数据写入内存缓存
        self.audio_buffer.extend(buffer)
        return len(buffer)

# 初始化配置与自定义流
speech_config = speechsdk.SpeechConfig(
    subscription=os.environ.get('SPEECH_KEY'),
    region=os.environ.get('SPEECH_REGION')
)
audio_callback = AudioBufferCallback()
pull_stream = speechsdk.audio.PullAudioOutputStream(audio_callback)
audio_config = speechsdk.audio.AudioOutputConfig(stream=pull_stream)

# 执行合成
speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config)
speech_synthesizer.speak_text_async("需要合成的文本").get()

# 合成完成后,audio_callback.audio_buffer中即为音频字节数据
print(f"获取到音频数据长度: {len(audio_callback.audio_buffer)} 字节")

内容的提问来源于stack exchange,提问作者jutta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 03:55:08