You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google TTS API生成MULAW 8KHz音频时遇编码错误求助

问题:Google文本转语音API生成MULAW 8KHz音频时触发编码错误

我尝试用Google文本转语音API生成MULAW 8KHz格式的音频,使用了官方提供的代码:

const text = 'Texte que vous souhaitez vocaliser'
const outputFile = 'testJaf.ulaw';
const languageCode = 'fr-FR';
const ssmlGender = 'FEMALE';
const sampleRateHertz = 8000;
const audioEncoding = "MULAW";

async function synthesizeWithEffectsProfile() {
    // Add one or more effects profiles to array.
    // Refer to documentation for more details:
    // https://cloud.google.com/text-to-speech/docs/audio-profiles
    //sampleRateHertz:sampleRateHertz,
    const effectsProfileId = ['telephony-class-application'];

    const request = {
      input: {text: text},
      voice: {languageCode: languageCode, ssmlGender: ssmlGender},
      audioConfig: {audioEncoding: audioEncoding, sampleRateHertz:sampleRateHertz, effectsProfileId: effectsProfileId},
    };

    const [response] = await speechClient.synthesizeSpeech(request);
    const writeFile = util.promisify(fs.writeFile);
    await writeFile(outputFile, response.audioContent, 'binary');
    console.log(`Audio content written to file: ${outputFile}`);
}

不管是否添加sampleRateHertz参数,只要使用MULAW格式就会触发编码类型错误(MP3和OGG格式可正常运行),错误信息如下:

(node:2256816) UnhandledPromiseRejectionWarning: Error: 3 INVALID_ARGUMENT: Invalid encoding type.
at Object.callErrorFromStatus (/root/node-red-vb/node_modules/@grpc/grpc-js/build/src/call.js:31:26)
at Object.onReceiveStatus (/root/node-red-vb/node_modules/@grpc/grpc-js/build/src/client.js:179:52)
at Object.onReceiveStatus (/root/node-red-vb/node_modules/@grpc/grpc-js/build/src/client-interceptors.js:336:141)
at Object.onReceiveStatus (/root/node-red-vb/node_modules/@grpc/grpc-js/build/src/client-interceptors.js:299:181)
at /root/node-red-vb/node_modules/@grpc/grpc-js/build/src/call-stream.js:145:78
at processTicksAndRejections (internal/process/task_queues.js:79:11)

请问是否需要配置额外参数才能使MULAW格式正常工作?


解答

问题出在参数组合的冲突上,调整后即可解决:

  1. 移除重复参数:telephony-class-application这个音效配置文件已经内置指定了MULAW编码+8000Hz采样率,不需要在audioConfig里重复声明sampleRateHertz,重复指定会导致参数冲突。
  2. 确认语音兼容性:部分语音类型不支持MULAW编码,建议切换为标准的Wavenet或神经语音(比如fr-FR-Wavenet-A)。

修正后的audioConfig可以简化为:

audioConfig: {
  audioEncoding: audioEncoding,
  effectsProfileId: effectsProfileId
}

如果仍报错,可直接移除音效配置,单独指定编码和采样率:

audioConfig: {
  audioEncoding: audioEncoding,
  sampleRateHertz: sampleRateHertz
}

这样就能正常生成符合要求的MULAW格式音频了。


内容的提问来源于stack exchange,提问作者jaf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 04:40:17