如何在Google Cloud Text-to-Speech中添加自定义时长停顿?
Google Cloud文本转语音添加停顿失效的解决办法
问题描述
我正在使用Google Cloud文本转语音模块,通过以下代码将文本转换为音频,但无法添加如5秒的停顿。我已在synthesis_input变量中添加了<break>标签,请问该如何解决?
import os from google.cloud import texttospeech os.environ["GOOGLE_APPLICATION_CREDENTIALS"]="G:\service-account-key.json" client = texttospeech.TextToSpeechClient() synthesis_input = texttospeech.SynthesisInput(text="<speak>You know that facebook is a place where millions of people share their thoughts. <break time=\"10s\"/> Today I am going to discuss 10 amazing things shared by people on facebook.</speak>") voice = texttospeech.VoiceSelectionParams(language_code='en-IN',name="en-IN-Wavenet-C",ssml_gender=texttospeech.SsmlVoiceGender.MALE) audio_config = texttospeech.AudioConfig( audio_encoding=texttospeech.AudioEncoding.MP3 ) response = client.synthesize_speech( input=synthesis_input, voice=voice, audio_config=audio_config ) with open("output.mp3", "wb") as out: # Write the response to the output file. out.write(response.audio_content) print('Audio content written to file "output.mp3"')
解决办法
问题出在你创建SynthesisInput时用了text参数——这个参数会把传入内容当作纯文本处理,直接忽略所有SSML标签。要让系统正确解析<break>这类语音控制标签,必须使用专门的ssml参数。
修改代码中创建synthesis_input的行,将text替换为ssml即可:
synthesis_input = texttospeech.SynthesisInput(ssml="<speak>You know that facebook is a place where millions of people share their thoughts. <break time=\"10s\"/> Today I am going to discuss 10 amazing things shared by people on facebook.</speak>")
你使用的en-IN-Wavenet-C语音模型原生支持SSML语法,修改后无需调整其他配置,系统就能生成指定时长的停顿。
内容的提问来源于stack exchange,提问作者Abhishek dot py
相关产品推荐
相关产品推荐

