You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google双向流式TTS响应中手动添加指定位置停顿?

解决Google Cloud流式TTS手动添加特定位置停顿的问题

核心结论

可以通过两种可靠方式实现特定位置的停顿:调整流式输入的SSML使用方式,或者利用请求间隔控制音频生成节奏。

方法一:正确使用SSML实现精准停顿

你之前尝试SSML无效,核心原因是流式请求未启用SSML模式。默认流式输入为纯文本模式,需在配置中明确指定使用SSML,且每个文本请求都用SSML格式包裹。

修改后的代码示例:

def run_streaming_tts_with_ssml_pauses():
    """Synthesizes speech with precise pauses using SSML in streaming mode."""
    from google.cloud import texttospeech
    import itertools

    client = texttospeech.TextToSpeechClient()

    # 启用SSML模式,配置语音与音频参数
    streaming_config = texttospeech.StreamingSynthesizeConfig(
        voice=texttospeech.VoiceSelectionParams(name="en-US-Journey-D", language_code="en-US"),
        audio_config=texttospeech.AudioConfig(audio_encoding=texttospeech.AudioEncoding.LINEAR16),
        # 关键:设置输入类型为SSML
        input_type=texttospeech.SynthesisInputType.SSML
    )

    config_request = texttospeech.StreamingSynthesizeRequest(streaming_config=streaming_config)

    def request_generator():
        # 使用<break>标签定义精准停顿,支持time(如200ms)或strength(如medium)参数
        yield texttospeech.StreamingSynthesizeRequest(
            input=texttospeech.StreamingSynthesisInput(ssml="<speak>Hello there.<break time='200ms'/></speak>")
        )
        yield texttospeech.StreamingSynthesizeRequest(
            input=texttospeech.StreamingSynthesisInput(ssml="<speak>How are you<break time='300ms'/></speak>")
        )
        yield texttospeech.StreamingSynthesizeRequest(
            input=texttospeech.StreamingSynthesisInput(ssml="<speak>today? It's<break time='150ms'/></speak>")
        )
        yield texttospeech.StreamingSynthesizeRequest(
            input=texttospeech.StreamingSynthesisInput(ssml="<speak>such nice weather outside.</speak>")
        )

    streaming_responses = client.streaming_synthesize(itertools.chain([config_request], request_generator()))
    for response in streaming_responses:
        print(f"Audio content size in bytes is: {len(response.audio_content)}")

注意事项:

  • 必须在StreamingSynthesizeConfig中设置input_type=texttospeech.SynthesisInputType.SSML
  • 所有SSML输入必须用<speak>根标签包裹,<break>标签可灵活控制停顿时长

方法二:利用请求间隔控制自然停顿

若不想使用SSML,可在生成流式请求时主动添加时间间隔,让TTS服务在处理完前一段文本后生成自然停顿。这种方式适合纯文本场景:

import time

def run_streaming_tts_with_pause_intervals():
    """Synthesizes speech with natural pauses using request intervals."""
    from google.cloud import texttospeech
    import itertools

    client = texttospeech.TextToSpeechClient()

    streaming_config = texttospeech.StreamingSynthesizeConfig(
        voice=texttospeech.VoiceSelectionParams(name="en-US-Journey-D", language_code="en-US")
    )

    config_request = texttospeech.StreamingSynthesizeRequest(streaming_config=streaming_config)

    def request_generator():
        yield texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text="Hello there. "))
        time.sleep(0.2)  # 添加200ms间隔模拟停顿
        yield texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text="How are you "))
        time.sleep(0.3)
        yield texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text="today? It's "))
        time.sleep(0.15)
        yield texttospeech.StreamingSynthesizeRequest(input=texttospeech.StreamingSynthesisInput(text="such nice weather outside."))

    streaming_responses = client.streaming_synthesize(itertools.chain([config_request], request_generator()))
    for response in streaming_responses:
        print(f"Audio content size in bytes is: {len(response.audio_content)}")

该方式无需修改输入格式,但停顿长度需根据实际效果调整,适合对精度要求不高的场景。

关于音频块与文本位置对齐的问题

Google Cloud流式TTS的响应确实未直接返回音频块与输入文本的映射关系,但可通过以下方式间接对齐:

  • 每个流式请求仅包含一段独立文本单元(如短语或句子),确保音频块(或连续小音频块)对应单一输入文本
  • 发送请求时记录时间戳,结合音频块接收时间,大致推断对应关系,方便后处理时添加静音

内容的提问来源于stack exchange,提问作者Adian Liusie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 08:27:38