You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Amazon Polly生成的音频添加逐词时间戳用于字幕制作?

解决Amazon Polly音频单词时间戳获取问题的方案

一、直接利用Amazon Polly内置功能获取时间戳

Amazon Polly原生支持返回单词级时间戳,只需在调用synthesize_speech API时指定输出格式为json并开启enable_word_time_stamps参数,返回的JSON数据中就会包含每个单词的开始、结束时间(单位:毫秒)。

示例Python代码:

import boto3
import json

# 初始化Polly客户端
polly_client = boto3.client('polly', region_name='us-east-1')

# 调用API生成带时间戳的音频数据
response = polly_client.synthesize_speech(
    Text="需要转换的目标文本",
    OutputFormat='json',
    VoiceId='Zhiyu',  # 根据需求选择对应语音ID
    EnableWordTimeStamps=True
)

# 解析时间戳信息
timestamps_data = json.loads(response['AudioStream'].read())
for word_item in timestamps_data['words']:
    print(f"单词: {word_item['word']}, 开始时间: {word_item['start_time']}ms, 结束时间: {word_item['end_time']}ms")

拿到时间戳后,可直接转换为SRT、VTT等标准字幕格式,用于视频制作。

二、替代方案(Polly功能受限场景)

如果无法使用Polly内置时间戳功能,可尝试以下两种方式:

1. 借助语音转文字(ASR)服务提取时间戳

将Polly生成的音频上传至Amazon Transcribe(或其他ASR服务),该服务可返回精确的单词级时间戳:

  • 把Polly生成的音频存储到S3存储桶
  • 调用Transcribe的start_transcription_job API,指定音频路径与输出格式
  • 任务完成后,从返回结果中提取每个单词的开始/结束时间

2. 开源工具本地处理音频

使用OpenAI的whisper或pyannote.audio等开源ASR模型,本地处理Polly生成的音频并提取时间戳。

示例Whisper代码:

import whisper

# 加载模型(可根据需求选择base/small/large等模型)
model = whisper.load_model("base")
# 转录音频并开启单词时间戳功能
result = model.transcribe("polly_generated_audio.mp3", word_timestamps=True)

# 遍历输出单词时间戳
for segment in result['segments']:
    for word in segment['words']:
        print(f"单词: {word['word']}, 开始时间: {word['start']}s, 结束时间: {word['end']}s")

该方案无需依赖云服务,适合对数据隐私要求较高的场景。

内容的提问来源于stack exchange,提问作者Hugo Novais

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 03:14:58