You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Google Cloud Speech-to-Text API生成细粒度时间分辨率SRT字幕

实现Google Cloud Speech-to-Text生成单/双单词细粒度SRT字幕

问题背景

现有Python代码通过Google Cloud Speech-to-Text API生成的SRT字幕是段落级的,每个条目包含多句文本,时间跨度较大:

1
00:00:00,040 --> 00:00:02,960
The sun set over the horizon
painting the sky and hues of

需求是生成细粒度SRT,每个条目仅包含1-2个单词,对应精确的单词级时间戳:

1
00:00:00,040 --> 00:00:00,540
The

2
00:00:00,540 --> 00:00:01,040
sun

实现方案

Google Cloud Speech-to-Text API默认生成的SRT是段落级的,但API已经返回了每个单词的时间戳(代码中已开启enable_word_time_offsets=True),因此可以通过解析单词级时间数据,自行生成细粒度SRT。

修改后的完整代码

from google.api_core.client_options import ClientOptions
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech
import json
from google.protobuf.json_format import MessageToDict

MAX_AUDIO_LENGTH_SECS = 8 * 60 * 60

def format_srt_time(seconds):
    """将秒数转换为SRT格式的时间字符串:HH:MM:SS,mmm"""
    hours = int(seconds // 3600)
    minutes = int((seconds % 3600) // 60)
    secs = seconds % 60
    milliseconds = int((secs - int(secs)) * 1000)
    return f"{hours:02d}:{minutes:02d}:{int(secs):02d},{milliseconds:03d}"

def generate_fine_grained_srt(words, max_words_per_entry=2):
    """根据单词列表生成细粒度SRT,每个条目最多包含max_words_per_entry个单词"""
    srt_lines = []
    entry_count = 1
    i = 0
    total_words = len(words)
    
    while i < total_words:
        # 确定当前条目的单词范围(1或2个)
        end_idx = min(i + max_words_per_entry, total_words)
        current_words = words[i:end_idx]
        
        # 获取条目的起始和结束时间
        start_time = current_words[0]['startTime']
        end_time = current_words[-1]['endTime']
        
        # 生成SRT条目
        srt_lines.append(str(entry_count))
        srt_lines.append(f"{format_srt_time(start_time)} --> {format_srt_time(end_time)}")
        srt_lines.append(' '.join([word['word'] for word in current_words]))
        srt_lines.append('')  # 空行分隔条目
        
        entry_count += 1
        i = end_idx
    
    return '\n'.join(srt_lines)

def run_batch_recognize():
    # 初始化客户端
    client = SpeechClient(
        client_options=ClientOptions(
            api_endpoint="us-central1-speech.googleapis.com",
        ),
    )

    # 音频文件的GCS地址
    audio_gcs_uri = "<redacted>"

    config = cloud_speech.RecognitionConfig(
        explicit_decoding_config=cloud_speech.ExplicitDecodingConfig(
            encoding=cloud_speech.ExplicitDecodingConfig.AudioEncoding.LINEAR16,
            sample_rate_hertz=24000,
            audio_channel_count=1,
        ),
        features=cloud_speech.RecognitionFeatures(
            enable_word_confidence=True,
            enable_word_time_offsets=True,
            enable_automatic_punctuation=True,
            max_alternatives=1,  # 只保留最优结果,减少处理量
        ),
        model="short",
        language_codes=["en-US"],
    )

    output_config = cloud_speech.RecognitionOutputConfig(
        inline_response_config=cloud_speech.InlineOutputConfig(),
    )

    files = [cloud_speech.BatchRecognizeFileMetadata(uri=audio_gcs_uri)]

    request = cloud_speech.BatchRecognizeRequest(
        recognizer="<redacted>",
        config=config,
        files=files,
        recognition_output_config=output_config,
    )
    operation = client.batch_recognize(request=request)

    print("等待识别完成...")
    response = operation.result(timeout=3 * MAX_AUDIO_LENGTH_SECS)

    # 转换为字典格式
    response_dict = MessageToDict(response._pb)
    
    # 提取单词级数据
    all_words = []
    for result in response_dict["results"][audio_gcs_uri]["inlineResult"]["transcriptResults"][0]["alternatives"][0]["words"]:
        # 解析时间(API返回的是字符串,比如"1.234s",需要转换为秒数)
        start_time = float(result["startTime"].replace("s", ""))
        end_time = float(result["endTime"].replace("s", ""))
        all_words.append({
            "word": result["word"],
            "startTime": start_time,
            "endTime": end_time
        })
    
    # 生成细粒度SRT
    fine_grained_srt = generate_fine_grained_srt(all_words, max_words_per_entry=2)
    
    # 输出结果
    print("细粒度SRT字幕:\n")
    print(fine_grained_srt)

run_batch_recognize()

关键修改说明

  1. 时间格式转换函数:format_srt_time将API返回的秒数转换为SRT标准的HH:MM:SS,mmm格式。
  2. 细粒度SRT生成函数:generate_fine_grained_srt遍历单词列表,按1-2个单词为一组生成SRT条目,每组的起始时间取第一个单词的开始时间,结束时间取最后一个单词的结束时间。
  3. 结果解析逻辑:不再直接使用API生成的段落级SRT,而是提取words字段中的单词级时间和文本数据,自行处理。
  4. 优化配置:将max_alternatives改为1,只处理最优识别结果,减少不必要的数据处理。

效果示例

运行修改后的代码,会生成如下格式的细粒度SRT:

1
00:00:00,040 --> 00:00:00,540
The

2
00:00:00,540 --> 00:00:01,040
sun

3
00:00:01,040 --> 00:00:01,540
set

内容的提问来源于stack exchange,提问作者theprogrammer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 13:17:08