You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置Google Speech-To-Text API生成.vtt/.srt字幕文件

Google Speech-To-Text API 配置输出为SRT/VTT格式

一、REST API调用时的参数配置

直接在请求体的config字段中添加outputFormat参数,指定为SRT或VTT即可。示例请求体如下:

{
  "config": {
    "encoding": "LINEAR16",
    "sampleRateHertz": 16000,
    "languageCode": "zh-CN",
    "outputFormat": "SRT"
  },
  "audio": {
    "uri": "gs://your-bucket/audio-file.wav"
  }
}

调用v1/speech:recognize(短音频)或v1/speech:longrunningrecognize(长音频)接口时,将上述请求体作为POST内容发送,返回结果即为对应格式的字幕文本。

二、客户端库(以Python为例)的配置

在创建RecognitionConfig对象时,指定output_format参数为对应的枚举值:

from google.cloud import speech_v1p1beta1 as speech

client = speech.SpeechClient()

# 配置音频来源
audio = speech.RecognitionAudio(uri="gs://your-bucket/audio-file.wav")

# 配置识别参数,指定输出格式
config = speech.RecognitionConfig(
    encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16,
    sample_rate_hertz=16000,
    language_code="zh-CN",
    output_format=speech.TranscriptOutputFormat.SRT
)

# 短音频识别
response = client.recognize(config=config, audio=audio)
# 长音频识别用以下代码
# operation = client.long_running_recognize(config=config, audio=audio)
# response = operation.result(timeout=90)

# 输出字幕内容
print(response.results[0].alternatives[0].transcript)

三、关键注意事项

  • 确保使用的API版本为v1或v1p1beta1,旧版本不支持outputFormat参数。
  • 处理时长超过1分钟的音频时,必须使用longrunningrecognize接口,配置方式与短音频一致。
  • 确认你的Google Cloud账号已启用Speech-To-Text API,并拥有对应的调用权限。

内容的提问来源于stack exchange,提问作者Paulo André de Andrade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 15:05:59