You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure认知服务Speech SDK Python:synthesizing回调音频杂音问题排查

通过synthesizing回调流式写入音频文件的杂音问题排查与解决

问题说明

需要通过synthesizing回调实现音频数据的流式写入文件,发现手动写入的server_bad_audio.wav存在杂音,但SDK自动写入的server_audio.wav播放正常,且必须使用synthesizing回调完成流式写入。

现有代码

audio_queue = asyncio.Queue()

async def send_audio(self, queue):
    with wave.open("server_bad_audio.wav", "wb") as wav_file:
        wav_file.setnchannels(1)
        wav_file.setsampwidth(SAMPLE_WIDTH)
        wav_file.setframerate(FRAME_RATE)
        while True:
            audio_data = await queue.get()
            if audio_data is None:
                break
            self.logger.info('Sending audio chunk of length {}'.format(len(audio_data)))
            wav_file.writeframes(audio_data)

def synthesize_callback(evt: SpeechSynthesisEventArgs):
    audio = evt.result.audio_data
    self.logger.info('Audio chunk received of length {}, duration {}'.format(len(audio), evt.result.audio_duration))
    audio_queue.put_nowait(audio)

# ... 省略其他代码
audio_config = AudioOutputConfig(filename="server_audio.wav")
synthesizer = SpeechSynthesizer(speech_config=self.speech_config, audio_config=audio_config)

synthesizer.synthesizing.connect(synthesize_callback)
result = synthesizer.speak_ssml_async(ssml_text).get()

# ... 省略其他代码
audio_queue.put_nowait(None)
await send_audio_task

问题根源

  1. 硬编码的音频格式参数不匹配
    你手动设置的WAV文件参数(单通道、SAMPLE_WIDTH、FRAME_RATE)和Speech SDK实际返回的audio_data格式不一致。SDK自动写入server_audio.wav时会使用合成音频的真实格式,所以播放正常;而手动写入时用了错误的格式参数,导致音频解码时出现杂音。

  2. 线程与异步队列的潜在顺序风险
    synthesizing回调运行在SDK的后台线程中,虽然put_nowait是线程安全的,但极端情况下可能出现数据块乱序,不过这不是主要杂音来源。

解决方法

方法1:动态获取音频格式参数

从回调的事件对象中获取真实的音频格式,再设置WAV文件参数,避免硬编码:

async def send_audio(self, queue, audio_format):
    with wave.open("server_fixed_audio.wav", "wb") as wav_file:
        wav_file.setnchannels(audio_format.channels)
        wav_file.setsampwidth(audio_format.bits_per_sample // 8)
        wav_file.setframerate(audio_format.samples_per_second)
        while True:
            audio_data = await queue.get()
            if audio_data is None:
                break
            self.logger.info('Sending audio chunk of length {}'.format(len(audio_data)))
            wav_file.writeframes(audio_data)

def synthesize_callback(evt: SpeechSynthesisEventArgs):
    audio = evt.result.audio_data
    # 首次回调时获取音频格式
    if not hasattr(self, 'audio_format'):
        self.audio_format = evt.result.audio_format
    self.logger.info('Audio chunk received of length {}, duration {}'.format(len(audio), evt.result.audio_duration))
    audio_queue.put_nowait(audio)

# 启动任务时传入获取到的音频格式
send_audio_task = asyncio.create_task(self.send_audio(audio_queue, self.audio_format))

方法2:避免同时启用SDK自动写入和回调

如果只需要通过回调流式处理音频,不要设置AudioOutputConfig(filename),改为使用无输出的配置,避免SDK内部的音频处理对回调数据产生干扰:

# 不写入文件,仅通过回调获取数据
audio_config = AudioOutputConfig(use_default_speaker=False)
# 或者直接传None
# audio_config = None
synthesizer = SpeechSynthesizer(speech_config=self.speech_config, audio_config=audio_config)

方法3:强化线程安全的队列操作

改用线程安全队列配合异步包装,确保数据不会丢失或乱序:

from queue import Queue as ThreadSafeQueue
import asyncio

audio_queue = ThreadSafeQueue()

async def send_audio(self, queue):
    with wave.open("server_fixed_audio.wav", "wb") as wav_file:
        # 动态设置音频格式参数
        wav_file.setnchannels(self.audio_format.channels)
        wav_file.setsampwidth(self.audio_format.bits_per_sample // 8)
        wav_file.setframerate(self.audio_format.samples_per_second)
        while True:
            audio_data = await asyncio.to_thread(queue.get)
            if audio_data is None:
                break
            self.logger.info('Sending audio chunk of length {}'.format(len(audio_data)))
            wav_file.writeframes(audio_data)

def synthesize_callback(evt: SpeechSynthesisEventArgs):
    audio = evt.result.audio_data
    if not hasattr(self, 'audio_format'):
        self.audio_format = evt.result.audio_format
    self.logger.info('Audio chunk received of length {}, duration {}'.format(len(audio), evt.result.audio_duration))
    audio_queue.put(audio)

内容的提问来源于stack exchange,提问作者Newbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 08:13:20