Azure认知服务Speech SDK Python:synthesizing回调音频杂音问题排查
通过synthesizing回调流式写入音频文件的杂音问题排查与解决
问题说明
需要通过synthesizing回调实现音频数据的流式写入文件,发现手动写入的server_bad_audio.wav存在杂音,但SDK自动写入的server_audio.wav播放正常,且必须使用synthesizing回调完成流式写入。
现有代码
audio_queue = asyncio.Queue() async def send_audio(self, queue): with wave.open("server_bad_audio.wav", "wb") as wav_file: wav_file.setnchannels(1) wav_file.setsampwidth(SAMPLE_WIDTH) wav_file.setframerate(FRAME_RATE) while True: audio_data = await queue.get() if audio_data is None: break self.logger.info('Sending audio chunk of length {}'.format(len(audio_data))) wav_file.writeframes(audio_data) def synthesize_callback(evt: SpeechSynthesisEventArgs): audio = evt.result.audio_data self.logger.info('Audio chunk received of length {}, duration {}'.format(len(audio), evt.result.audio_duration)) audio_queue.put_nowait(audio) # ... 省略其他代码 audio_config = AudioOutputConfig(filename="server_audio.wav") synthesizer = SpeechSynthesizer(speech_config=self.speech_config, audio_config=audio_config) synthesizer.synthesizing.connect(synthesize_callback) result = synthesizer.speak_ssml_async(ssml_text).get() # ... 省略其他代码 audio_queue.put_nowait(None) await send_audio_task
问题根源
硬编码的音频格式参数不匹配
你手动设置的WAV文件参数(单通道、SAMPLE_WIDTH、FRAME_RATE)和Speech SDK实际返回的audio_data格式不一致。SDK自动写入server_audio.wav时会使用合成音频的真实格式,所以播放正常;而手动写入时用了错误的格式参数,导致音频解码时出现杂音。线程与异步队列的潜在顺序风险
synthesizing回调运行在SDK的后台线程中,虽然put_nowait是线程安全的,但极端情况下可能出现数据块乱序,不过这不是主要杂音来源。
解决方法
方法1:动态获取音频格式参数
从回调的事件对象中获取真实的音频格式,再设置WAV文件参数,避免硬编码:
async def send_audio(self, queue, audio_format): with wave.open("server_fixed_audio.wav", "wb") as wav_file: wav_file.setnchannels(audio_format.channels) wav_file.setsampwidth(audio_format.bits_per_sample // 8) wav_file.setframerate(audio_format.samples_per_second) while True: audio_data = await queue.get() if audio_data is None: break self.logger.info('Sending audio chunk of length {}'.format(len(audio_data))) wav_file.writeframes(audio_data) def synthesize_callback(evt: SpeechSynthesisEventArgs): audio = evt.result.audio_data # 首次回调时获取音频格式 if not hasattr(self, 'audio_format'): self.audio_format = evt.result.audio_format self.logger.info('Audio chunk received of length {}, duration {}'.format(len(audio), evt.result.audio_duration)) audio_queue.put_nowait(audio) # 启动任务时传入获取到的音频格式 send_audio_task = asyncio.create_task(self.send_audio(audio_queue, self.audio_format))
方法2:避免同时启用SDK自动写入和回调
如果只需要通过回调流式处理音频,不要设置AudioOutputConfig(filename),改为使用无输出的配置,避免SDK内部的音频处理对回调数据产生干扰:
# 不写入文件,仅通过回调获取数据 audio_config = AudioOutputConfig(use_default_speaker=False) # 或者直接传None # audio_config = None synthesizer = SpeechSynthesizer(speech_config=self.speech_config, audio_config=audio_config)
方法3:强化线程安全的队列操作
改用线程安全队列配合异步包装,确保数据不会丢失或乱序:
from queue import Queue as ThreadSafeQueue import asyncio audio_queue = ThreadSafeQueue() async def send_audio(self, queue): with wave.open("server_fixed_audio.wav", "wb") as wav_file: # 动态设置音频格式参数 wav_file.setnchannels(self.audio_format.channels) wav_file.setsampwidth(self.audio_format.bits_per_sample // 8) wav_file.setframerate(self.audio_format.samples_per_second) while True: audio_data = await asyncio.to_thread(queue.get) if audio_data is None: break self.logger.info('Sending audio chunk of length {}'.format(len(audio_data))) wav_file.writeframes(audio_data) def synthesize_callback(evt: SpeechSynthesisEventArgs): audio = evt.result.audio_data if not hasattr(self, 'audio_format'): self.audio_format = evt.result.audio_format self.logger.info('Audio chunk received of length {}, duration {}'.format(len(audio), evt.result.audio_duration)) audio_queue.put(audio)
内容的提问来源于stack exchange,提问作者Newbie
相关产品推荐
相关产品推荐

