如何禁用Azure TTS生成音频样本时的自动播放功能?
解决Azure Speech Service TTS禁用音频播放的问题
问题原因
直接初始化AudioOutputConfig()时,SDK没有明确的输出目标,会触发ValueError;而显式设置use_default_speaker=False时,SDK仍会默认调用系统音频输出设备,导致音频播放。
正确解决方案
使用SDK提供的NullAudioOutputStream,它会直接丢弃合成的音频数据,既不播放也无需写入文件。代码示例如下:
import os import azure.cognitiveservices.speech as speechsdk speech_config = speechsdk.SpeechConfig( subscription=os.environ.get('SPEECH_KEY'), region=os.environ.get('SPEECH_REGION') ) # 创建空音频输出流,直接丢弃合成的音频数据 null_stream = speechsdk.audio.NullAudioOutputStream() audio_config = speechsdk.audio.AudioOutputConfig(stream=null_stream) # 初始化合成器并执行文本合成 speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config) result = speech_synthesizer.speak_text_async("需要合成的目标文本").get() # 可选:检查合成结果状态 if result.reason == speechsdk.ResultReason.SynthesizingAudioCompleted: print("文本合成完成") elif result.reason == speechsdk.ResultReason.Canceled: cancellation_details = result.cancellation_details print(f"合成任务取消: {cancellation_details.reason}") if cancellation_details.reason == speechsdk.CancellationReason.Error: print(f"错误详情: {cancellation_details.error_details}")
补充:获取音频数据但不播放
如果需要在不播放的前提下获取合成的音频字节数据,可以自定义PullAudioOutputStream缓存数据到内存:
import os import azure.cognitiveservices.speech as speechsdk # 自定义拉取式音频流回调,用于缓存音频数据 class AudioBufferCallback(speechsdk.audio.PullAudioOutputStreamCallback): def __init__(self): super().__init__() self.audio_buffer = bytearray() def pull(self, buffer: memoryview) -> int: # 将音频数据写入内存缓存 self.audio_buffer.extend(buffer) return len(buffer) # 初始化配置与自定义流 speech_config = speechsdk.SpeechConfig( subscription=os.environ.get('SPEECH_KEY'), region=os.environ.get('SPEECH_REGION') ) audio_callback = AudioBufferCallback() pull_stream = speechsdk.audio.PullAudioOutputStream(audio_callback) audio_config = speechsdk.audio.AudioOutputConfig(stream=pull_stream) # 执行合成 speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config) speech_synthesizer.speak_text_async("需要合成的文本").get() # 合成完成后,audio_callback.audio_buffer中即为音频字节数据 print(f"获取到音频数据长度: {len(audio_callback.audio_buffer)} 字节")
内容的提问来源于stack exchange,提问作者jutta
相关产品推荐
相关产品推荐

