如何在Google双向流式TTS响应中手动添加指定位置停顿?
解决Google Cloud流式TTS手动添加特定位置停顿的问题
核心结论
可以通过两种可靠方式实现特定位置的停顿:调整流式输入的SSML使用方式,或者利用请求间隔控制音频生成节奏。
方法一:正确使用SSML实现精准停顿
你之前尝试SSML无效,核心原因是流式请求未启用SSML模式。默认流式输入为纯文本模式,需在配置中明确指定使用SSML,且每个文本请求都用SSML格式包裹。
修改后的代码示例:
def run_streaming_tts_with_ssml_pauses(): """Synthesizes speech with precise pauses using SSML in streaming mode.""" from google.cloud import texttospeech import itertools client = texttospeech.TextToSpeechClient() # 启用SSML模式,配置语音与音频参数 streaming_config = texttospeech.StreamingSynthesizeConfig( voice=texttospeech.VoiceSelectionParams(name="en-US-Journey-D", language_code="en-US"), audio_config=texttospeech.AudioConfig(audio_encoding=texttospeech.AudioEncoding.LINEAR16), # 关键:设置输入类型为SSML input_type=texttospeech.SynthesisInputType.SSML ) config_request = texttospeech.StreamingSynthesizeRequest(streaming_config=streaming_config) def request_generator(): # 使用<break>标签定义精准停顿,支持time(如200ms)或strength(如medium)参数 yield texttospeech.StreamingSynthesizeRequest( input=texttospeech.StreamingSynthesisInput(ssml="<speak>Hello there.<break time='200ms'/></speak>") ) yield texttospeech.StreamingSynthesizeRequest( input=texttospeech.StreamingSynthesisInput(ssml="<speak>How are you<break time='300ms'/></speak>") ) yield texttospeech.StreamingSynthesizeRequest( input=texttospeech.StreamingSynthesisInput(ssml="<speak>today? It's<break time='150ms'/></speak>") ) yield texttospeech.StreamingSynthesizeRequest( input=texttospeech.StreamingSynthesisInput(ssml="<speak>such nice weather outside.</speak>") ) streaming_responses = client.streaming_synthesize(itertools.chain([config_request], request_generator())) for response in streaming_responses: print(f"Audio content size in bytes is: {len(response.audio_content)}")
注意事项:
- 必须在
StreamingSynthesizeConfig中设置input_type=texttospeech.SynthesisInputType.SSML - 所有SSML输入必须用
<speak>根标签包裹,<break>标签可灵活控制停顿时长
方法二:利用请求间隔控制自然停顿
若不想使用SSML,可在生成流式请求时主动添加时间间隔,让TTS服务在处理完前一段文本后生成自然停顿。这种方式适合纯文本场景:
import time def run_streaming_tts_with_pause_intervals(): """Synthesizes speech with natural pauses using request intervals.""" from google.cloud import texttospeech import itertools client = texttospeech.TextToSpeechClient() streaming_config = texttospeech.StreamingSynthesizeConfig( voice=texttospeech.VoiceSelectionParams(name="en-US-Journey-D", language_code="en-US") ) config_request = texttospeech.StreamingSynthesizeRequest(streaming_config=streaming_config) def request_generator(): yield texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text="Hello there. ")) time.sleep(0.2) # 添加200ms间隔模拟停顿 yield texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text="How are you ")) time.sleep(0.3) yield texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text="today? It's ")) time.sleep(0.15) yield texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text="such nice weather outside.")) streaming_responses = client.streaming_synthesize(itertools.chain([config_request], request_generator())) for response in streaming_responses: print(f"Audio content size in bytes is: {len(response.audio_content)}")
该方式无需修改输入格式,但停顿长度需根据实际效果调整,适合对精度要求不高的场景。
关于音频块与文本位置对齐的问题
Google Cloud流式TTS的响应确实未直接返回音频块与输入文本的映射关系,但可通过以下方式间接对齐:
- 每个流式请求仅包含一段独立文本单元(如短语或句子),确保音频块(或连续小音频块)对应单一输入文本
- 发送请求时记录时间戳,结合音频块接收时间,大致推断对应关系,方便后处理时添加静音
内容的提问来源于stack exchange,提问作者Adian Liusie
相关产品推荐
相关产品推荐

