You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Cloud Text-to-Speech中添加自定义时长停顿?

Google Cloud文本转语音添加停顿失效的解决办法

问题描述

我正在使用Google Cloud文本转语音模块,通过以下代码将文本转换为音频,但无法添加如5秒的停顿。我已在synthesis_input变量中添加了<break>标签,请问该如何解决?

import os
from google.cloud import texttospeech

os.environ["GOOGLE_APPLICATION_CREDENTIALS"]="G:\service-account-key.json"

client = texttospeech.TextToSpeechClient()

synthesis_input = texttospeech.SynthesisInput(text="<speak>You know that facebook is a place where millions of people share their thoughts. <break time=\"10s\"/> Today I am going to discuss 10 amazing things shared by people on facebook.</speak>")
voice = texttospeech.VoiceSelectionParams(language_code='en-IN',name="en-IN-Wavenet-C",ssml_gender=texttospeech.SsmlVoiceGender.MALE)

audio_config = texttospeech.AudioConfig(
audio_encoding=texttospeech.AudioEncoding.MP3
)
response = client.synthesize_speech(
input=synthesis_input, voice=voice, audio_config=audio_config
)
with open("output.mp3", "wb") as out:
# Write the response to the output file.
out.write(response.audio_content)
print('Audio content written to file "output.mp3"')

解决办法

问题出在你创建SynthesisInput时用了text参数——这个参数会把传入内容当作纯文本处理,直接忽略所有SSML标签。要让系统正确解析<break>这类语音控制标签,必须使用专门的ssml参数。

修改代码中创建synthesis_input的行,将text替换为ssml即可:

synthesis_input = texttospeech.SynthesisInput(ssml="<speak>You know that facebook is a place where millions of people share their thoughts. <break time=\"10s\"/> Today I am going to discuss 10 amazing things shared by people on facebook.</speak>")

你使用的en-IN-Wavenet-C语音模型原生支持SSML语法,修改后无需调整其他配置,系统就能生成指定时长的停顿。

内容的提问来源于stack exchange,提问作者Abhishek dot py

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 03:55:17