如何通过SSML优化Azure TTS en-US-NancyNeural的低语效果?
调整Azure TTS en-US-NancyNeural低语风格的SSML方案与避坑指南
背景
我正在开展视频配音项目,选用Azure TTS的免费额度(截至2024年1月23日每月50万字符),选中en-US-NancyNeural语音的低语风格,但默认效果语速偏快、音量偏响,需要通过SSML调整得更轻柔、缓慢、自然。
有效调整方案
通过SSML的<prosody>标签配合<mstts:express-as>,可以精准控制语音的语速、音量和语调,适配低语场景:
- 语速控制:使用
rate参数调整,推荐设置为75%-85%,既保证缓慢自然,又不会出现发音卡顿。 - 音量控制:使用
volume参数降低音量,推荐设置为**-8dB到-12dB**,贴合低语的轻柔感。 - 语调微调:可选
pitch参数,设置为**-5%到-10%**,让语气更平缓,避免尖锐感。
修改后的核心SSML结构如下:
<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xmlns:mstts="https://www.w3.org/2001/mstts" xml:lang="en-US"> <voice name="en-US-NancyNeural"> <mstts:express-as style="whispering"> <prosody rate="80%" volume="-10dB" pitch="-5%"> ${text} </prosody> </mstts:express-as> </voice> </speak>
完整NodeJS实现代码
async function generateSpeechFromText(name, text, tempDirectory) { console.log(`Generating speech from text for section: ${name}`) const audioFile = `${tempDirectory}/${name}.wav` const speechConfig = TTSSdk.SpeechConfig.fromSubscription( process.env.AZURE_TTS_KEY, process.env.AZURE_TTS_REGION ) const audioConfig = TTSSdk.AudioConfig.fromAudioFileOutput(audioFile) speechConfig.speechSynthesisVoiceName = "en-US-NancyNeural" let synthesizer = new TTSSdk.SpeechSynthesizer(speechConfig, audioConfig) const ssml = `<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xmlns:mstts="https://www.w3.org/2001/mstts" xml:lang="en-US"> <voice name="${speechConfig.speechSynthesisVoiceName}"> <mstts:express-as style="whispering"> <prosody rate="80%" volume="-10dB" pitch="-5%"> ${text} </prosody> </mstts:express-as> </voice> </speak>` return new Promise((resolve, reject) => { synthesizer.speakSsmlAsync( ssml, (result) => { if (result.reason === TTSSdk.ResultReason.SynthesizingAudioCompleted) { console.log("Synthesis finished for: " + name) resolve(audioFile) } else { console.error( "Speech synthesis failed for: " + name, result.errorDetails ) reject(result.errorDetails) } synthesizer.close() }, (err) => { console.error("Error during synthesis for: " + name, err) synthesizer.close() reject(err) } ) }) }
避坑指南
- 标签嵌套顺序不能乱:必须按照
<voice>→<mstts:express-as>→<prosody>的顺序嵌套,若将<prosody>放在<mstts:express-as>外层,低语风格会被覆盖失效。 - 参数不要过度调整:语速低于70%会导致发音断裂卡顿,音量低于-15dB会几乎听不清,语调调整幅度过大(±20%以上)会让语音变得怪异。
- 先测试小片段再批量生成:先拿1-2句文本测试参数效果,确认符合预期后再批量处理,避免浪费免费额度和时间。
- 确保命名空间正确:SSML头部的
xmlns和xmlns:mstts属性必须准确,否则自定义标签不会生效。
内容的提问来源于stack exchange,提问作者Simon Nazarenko
相关产品推荐
相关产品推荐

