You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python调用Google Cloud Text-to-Speech插入外部音频遇错误求助

Google Cloud Text-to-Speech插入外部音频触发500内部错误

输入文本版本说明

  • 版本1(正常运行)
<speak version="1.1"
        xmlns="http://www.w3.org/2001/10/synthesis"
        xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
        xsi:schemaLocation="http://www.w3.org/2001/10/synthesis
                  http://www.w3.org/TR/speech-synthesis11/synthesis.xsd"
        xml:lang="en-GB">

The rain in Spain stays mainly in the plain.

How kind of you to let me come.

</speak>
  • 版本2(插入外部音频后触发错误)
    与版本1内容一致,仅添加<audio>标签引入外部音频:
<speak version="1.1"
        xmlns="http://www.w3.org/2001/10/synthesis"
        xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
        xsi:schemaLocation="http://www.w3.org/2001/10/synthesis
                  http://www.w3.org/TR/speech-synthesis11/synthesis.xsd"
        xml:lang="en-GB">

The rain in Spain stays mainly in the plain.
<audio src="uhm_male.mp3" />
How kind of you to let me come.

</speak>
  • 版本3(简化<speak>标签,仍触发相同错误)
    仅简化<speak>标签结构,其余与版本2一致,错误未消失:
<speak>

The rain in Spain stays mainly in the plain.
<audio src="uhm_male.mp3" />
How kind of you to let me come.

</speak>
  • 版本4(移除<audio>标签,恢复正常运行)
    与版本3内容一致,仅移除<audio>标签,可正常生成音频。

调用的Python代码

import os
from google.cloud import texttospeech_v1

os.environ['GOOGLE_APPLICATION_CREDENTIALS'] =\
 'not_my_real_credentials.json'


def getText(infile_name):
    with open(infile_name, 'r') as fobj:
        intext = fobj.read()
    return intext    

def setVoice(language_code, name, ssml_gender):
    theVoice = tts.VoiceSelectionParams(
        language_code=language_code,
        name=name,
        ssml_gender=ssml_gender
    )
    return theVoice

def doAudioConfig(speaking_rate, pitch, volume_gain_db):
    audioConfig = tts.AudioConfig(
        audio_encoding = tts.AudioEncoding.MP3,
        speaking_rate = speaking_rate,
        pitch = pitch,
        volume_gain_db = volume_gain_db
    )
    return audioConfig

# 依次测试以下文件:

# 正常运行
infile_name = './texts/test1.txt'

# 触发错误
#infile_name = './texts/test2.txt'

# 触发相同错误
#infile_name = './texts/test3.txt'

# 正常运行
infile_name = './texts/test4.txt'

outfile_name = './audio/audioOutput.mp3'

tts = texttospeech_v1
client = tts.TextToSpeechClient()
language_code = "en-GB"
name = "en-GB-Wavenet-F"
ssml_gender = "FEMALE"
pitch = -8.0
speaking_rate = 0.9
volume_gain_db = 0

intext = getText(infile_name)
print(f'\n{intext}\n')

theVoice = setVoice(language_code, name, ssml_gender)
audioConfig = doAudioConfig(speaking_rate, pitch, volume_gain_db)
synthesis_input = tts.SynthesisInput(ssml=intext)

response = client.synthesize_speech(
input=synthesis_input, voice=theVoice, audio_config=audioConfig)

with open(outfile_name, 'wb') as output1:
    output1.write(response.audio_content)

触发的错误信息

使用版本2或3的文本时,会抛出500内部服务器错误,具体栈信息如下:

Traceback (most recent call last):
  File "D:\py\_new\envo\lib\site-packages\google\api_core\grpc_helpers.py", line 66, in error_remapped_callable
    return callable_(*args, **kwargs)
  File "D:\py\_new\envo\lib\site-packages\grpc\_channel.py", line 946, in __call__
    return _end_unary_response_blocking(state, call, False, None)
  File "D:\py\_new\envo\lib\site-packages\grpc\_channel.py", line 849, in _end_unary_response_blocking
    raise _InactiveRpcError(state)
grpc._channel._InactiveRpcError: <_InactiveRpcError of RPC that terminated with:
        status = StatusCode.INTERNAL
        details = "Internal error encountered."
        debug_error_string = "{"created":"@1681519238.596000000","description":"Error received from peer ipv4:142.250.70.170:443","file":"src/core/lib/surface/call.cc","file_line":1075,"grpc_message":"Internal error encountered.","grpc_status":13}"
>

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "D:\py\_new\ttsgcp2\insertAudio.py", line 62, in <module>
    response = client.synthesize_speech(
  File "D:\py\_new\envo\lib\site-packages\google\cloud\texttospeech_v1\services\text_to_speech\client.py", line 497, in synthesize_speech
    response = rpc(request, retry=retry, timeout=timeout, metadata=metadata,)
  File "D:\py\_new\envo\lib\site-packages\google\api_core\gapic_v1\method.py", line 154, in __call__
    return wrapped_func(*args, **kwargs)
  File "D:\py\_new\envo\lib\site-packages\google\api_core\grpc_helpers.py", line 68, in error_remapped_callable
    raise exceptions.from_grpc_error(exc) from exc
google.api_core.exceptions.InternalServerError: 500 Internal error encountered.

已尝试的排查操作

  • 将音频文件格式更换为.wav和.ogg,错误依旧出现
  • 替换导入texttospeech而非texttospeech_v1,问题未得到改善
  • 待插入的外部音频为0.302秒的“uhm”音效,本地可正常播放

内容的提问来源于stack exchange,提问作者user3567761

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 20:17:00