You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure TTS服务中SSML音调提升无效问题求助

Azure TTS SSML音调调整无效的问题排查与修复

问题根源分析

你的代码里存在三个核心问题导致SSML的音调调整不生效:

  1. SSML未被实际传入合成函数
    在get_audio_and_text中,调用generate_ssml(response)生成了SSML文本,但没有将返回值传递给get_output_audio_file,仍然传入原始的纯文本response,导致SSML完全没被使用。
  2. SSML标签语法错误
    generate_ssml函数生成的SSML存在标签不匹配:出现了</voice>结束标签,但没有对应的<voice>开始标签,这种无效的SSML会被Azure TTS服务直接忽略,降级为纯文本处理。
  3. 使用了错误的合成方法
    get_output_audio_file里调用的speak_text_async是纯文本合成方法,无法解析SSML结构,必须使用专门处理SSML的speak_ssml_async方法。

修复后的关键代码片段

1. 修正SSML生成函数

def generate_ssml(response):
    # 修正标签结构,移除多余的</voice>,如需指定语音可添加<voice>标签
    ssml_text = f'<speak><prosody pitch="+15.00%">{response}</prosody></speak>'
    # 若需指定语音,可改为以下格式:
    # ssml_text = f'<speak><voice name="{get_value_from_json_key("voice-name")}"><prosody pitch="+15.00%">{response}</prosody></voice></speak>'
    return ssml_text

2. 传递SSML到合成函数

def get_audio_and_text(message_content, message_author):
    response = generate_conversation(message_content, message_author)
    ssml_response = generate_ssml(response)  # 接收生成的SSML
    get_output_audio_file(ssml_response)  # 传入SSML而非原始文本
    audio_file = output_file_name_with_path
    time.sleep(2)
    open('output.txt', 'w').close()
    print('------------------------------------------------------')
    os.remove(audio_file)

3. 使用SSML专用合成方法

def get_output_audio_file(ssml_text):
    speech_config = speechsdk.SpeechConfig(subscription=get_value_from_json_key("microsoft-azure-api-key"),
                                           region=get_value_from_json_key("microsoft-azure-speech-region"))
    speech_config.speech_synthesis_voice_name = get_value_from_json_key("voice-name")
    audio_config = speechsdk.audio.AudioOutputConfig(use_default_speaker=True)
    speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config)
      
    print("<Speaking...>")
    with open("output.txt", "a", encoding="utf-8") as out:
        out.write(str(ssml_text) + "\n")
    # 替换为SSML专用合成方法
    speech_synthesis_result = speech_synthesizer.speak_ssml_async(ssml_text).get()
    get_audio_or_return_error(speech_synthesis_result)

内容的提问来源于stack exchange,提问作者cwp0627

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 04:28:11