Azure TTS服务中SSML音调提升无效问题求助
Azure TTS SSML音调调整无效的问题排查与修复
问题根源分析
你的代码里存在三个核心问题导致SSML的音调调整不生效:
- SSML未被实际传入合成函数
在get_audio_and_text中,调用generate_ssml(response)生成了SSML文本,但没有将返回值传递给get_output_audio_file,仍然传入原始的纯文本response,导致SSML完全没被使用。 - SSML标签语法错误
generate_ssml函数生成的SSML存在标签不匹配:出现了</voice>结束标签,但没有对应的<voice>开始标签,这种无效的SSML会被Azure TTS服务直接忽略,降级为纯文本处理。 - 使用了错误的合成方法
get_output_audio_file里调用的speak_text_async是纯文本合成方法,无法解析SSML结构,必须使用专门处理SSML的speak_ssml_async方法。
修复后的关键代码片段
1. 修正SSML生成函数
def generate_ssml(response): # 修正标签结构,移除多余的</voice>,如需指定语音可添加<voice>标签 ssml_text = f'<speak><prosody pitch="+15.00%">{response}</prosody></speak>' # 若需指定语音,可改为以下格式: # ssml_text = f'<speak><voice name="{get_value_from_json_key("voice-name")}"><prosody pitch="+15.00%">{response}</prosody></voice></speak>' return ssml_text
2. 传递SSML到合成函数
def get_audio_and_text(message_content, message_author): response = generate_conversation(message_content, message_author) ssml_response = generate_ssml(response) # 接收生成的SSML get_output_audio_file(ssml_response) # 传入SSML而非原始文本 audio_file = output_file_name_with_path time.sleep(2) open('output.txt', 'w').close() print('------------------------------------------------------') os.remove(audio_file)
3. 使用SSML专用合成方法
def get_output_audio_file(ssml_text): speech_config = speechsdk.SpeechConfig(subscription=get_value_from_json_key("microsoft-azure-api-key"), region=get_value_from_json_key("microsoft-azure-speech-region")) speech_config.speech_synthesis_voice_name = get_value_from_json_key("voice-name") audio_config = speechsdk.audio.AudioOutputConfig(use_default_speaker=True) speech_synthesizer = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_config) print("<Speaking...>") with open("output.txt", "a", encoding="utf-8") as out: out.write(str(ssml_text) + "\n") # 替换为SSML专用合成方法 speech_synthesis_result = speech_synthesizer.speak_ssml_async(ssml_text).get() get_audio_or_return_error(speech_synthesis_result)
内容的提问来源于stack exchange,提问作者cwp0627
相关产品推荐
相关产品推荐

