You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google TTS SSML生成音频出现高频哨声问题求助

问题描述

使用Google Text-to-Speech生成音频时,输入的SSML内容如下:

<break time="6s"/>AIRDASS<break time="1s"/> introduction.<break time="9s"/>In this lesson, the following topics will be covered:<break time="1s"/> introduction of the AIRDASS system, <break time="1s"/>description of the main characteristics, <break time="1s"/>description of the main functionalities.

生成的音频在本该静音的<break>位置出现恼人的高频哨声。当前使用的配置参数如下:

{
    audioConfig: {
        audioEncoding: "LINEAR16",
        pitch: 0,
        speakingRate: 1
    },
    input: {
        ssml
    },
    voice: {
        languageCode: "en-GB",
        name: "en-GB-Neural2-D"
    }
}

请问是否有其他可添加的设置来消除该噪音?

解决方案建议
  • 更换音频编码格式:LINEAR16为未压缩PCM格式,易暴露底层音频瑕疵。可切换为MP3或OGG_OPUS压缩格式,这类格式通常会自动过滤高频噪音。修改配置示例:

    audioConfig: {
        audioEncoding: "MP3", // 或 "OGG_OPUS"
        pitch: 0,
        speakingRate: 1
    }
    
  • 调整SSML静音标签:避免使用过长的<break>(如6s、9s),拆分短分段静音或尝试<silence>标签。例如将<break time="6s"/>替换为6个<break time="1s"/>,或使用:

    <silence strength="none" time="6s"/>
    

    注意:部分Google TTS语音模型对<silence>标签支持有差异,需测试验证。

  • 更换语音模型:en-GB-Neural2-D可能存在特定音频瑕疵,尝试同语言其他Neural2模型(如en-GB-Neural2-C)或WaveNet模型(en-GB-Wavenet-D),不同模型的生成逻辑差异可避免哨声问题。

  • 添加音效配置参数:在audioConfig中加入effectsProfileId,指定场景优化配置(含降噪处理)。示例:

    audioConfig: {
        audioEncoding: "LINEAR16",
        pitch: 0,
        speakingRate: 1,
        effectsProfileId: ["telephony-class-application"]
    }
    

内容的提问来源于stack exchange,提问作者lviggiani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 10:28:19