You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Viseme数据在Linux系统中无法完整生成的原因排查

问题分析与修复方案

核心问题

你的代码在原生Ubuntu 22.04环境下Viseme数据不完整,主要是同步阻塞与异步事件未正确等待导致的,具体原因:

  1. 同步调用阻塞回调执行:future.get()是同步阻塞方法,会卡住当前线程。WSL环境线程调度机制特殊,回调能在阻塞期间完成;但原生Linux环境中,后台Viseme回调线程还没处理完所有事件,主线程就已拿到合成结果并处理viseme_data,导致只捕获了部分数据。
  2. 未实际等待Viseme完成事件:你创建了wait_for_visemes()异步任务但未等待它完成,直接调用send_audio_and_viseme,此时viseme_received_event还未触发,viseme_data仅收集到开头内容。
  3. Synthesizer对象可能被提前回收:获取合成结果后,synthesizer可能被Python垃圾回收机制销毁,导致后续Viseme回调无法触发。

修复后的代码

async def generate_speech(self, text, websocket, language="en", voice="female"):
    text = self.clean_text(text)
    speaker = self.voices.get(language, {}).get(voice, "en-US-AriaNeural")
    self.speech_config.speech_synthesis_voice_name = speaker

    viseme_data = []
    viseme_received_event = asyncio.Event()
    synthesis_completed_event = asyncio.Event()

    def viseme_callback(evt):
        viseme_info = {
            "timestamp": evt.audio_offset / 10000,
            "viseme_id": evt.viseme_id
        }
        viseme_data.append(viseme_info)
        if evt.viseme_id == speechsdk.VisemeId.EndOfSentence: 
            viseme_received_event.set()

    def synthesis_completed_callback(evt):
        if evt.result.reason == speechsdk.ResultReason.SynthesizingAudioCompleted:
            synthesis_completed_event.set()
        elif evt.result.reason == speechsdk.ResultReason.Canceled:
            cancellation_details = evt.result.cancellation_details
            print(f"Speech synthesis failed: {cancellation_details.reason}")
            if cancellation_details.reason == speechsdk.CancellationReason.Error:
                print(f"Error Details: {cancellation_details.error_details}")
            synthesis_completed_event.set()
            raise Exception(f"Speech synthesis failed: {cancellation_details.reason}")

    # 创建synthesizer并绑定回调
    synthesizer = speechsdk.SpeechSynthesizer(speech_config=self.speech_config)
    synthesizer.viseme_received.connect(viseme_callback)
    synthesizer.synthesis_completed.connect(synthesis_completed_callback)

    ssml_template = f"""
    <speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xmlns:mstts="http://www.w3.org/2001/mstts" xml:lang="{language}">
        <voice name="{speaker}">
            <mstts:viseme type="facialExpression"/>
            {text}
        </voice>
    </speak>
    """

    try:
        # 异步启动合成,不阻塞主线程
        synthesizer.speak_ssml_async(ssml_template)
        
        # 等待合成完成和Viseme事件全部接收
        await asyncio.gather(
            synthesis_completed_event.wait(),
            viseme_received_event.wait()
        )

        # 确保synthesizer在回调完成后再被回收
        audio_stream = io.BytesIO(synthesizer.get_last_result().audio_data)
        await self.send_audio_and_viseme(audio_stream, websocket, viseme_data)

    except Exception as e:
        print(f"Error during speech synthesis: {str(e)}")
        raise
    finally:
        # 显式释放资源
        synthesizer.close()

关键修改点

  • 新增synthesis_completed_event,通过synthesis_completed回调监听合成完成状态,避免同步阻塞获取结果。
  • 使用asyncio.gather()同时等待合成完成和Viseme结束事件,确保所有数据收集完毕后再处理。
  • 显式保留synthesizer引用直到所有事件完成,避免被提前回收;最后调用close()释放资源。
  • 替换同步的future.get()为异步事件等待,让主线程不阻塞,保证回调线程能完整处理所有Viseme数据。

内容的提问来源于stack exchange,提问作者Abstract

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 07:04:51