如何通过Google Cloud Speech-to-Text API生成细粒度时间分辨率SRT字幕
实现Google Cloud Speech-to-Text生成单/双单词细粒度SRT字幕
问题背景
现有Python代码通过Google Cloud Speech-to-Text API生成的SRT字幕是段落级的,每个条目包含多句文本,时间跨度较大:
1 00:00:00,040 --> 00:00:02,960 The sun set over the horizon painting the sky and hues of
需求是生成细粒度SRT,每个条目仅包含1-2个单词,对应精确的单词级时间戳:
1 00:00:00,040 --> 00:00:00,540 The 2 00:00:00,540 --> 00:00:01,040 sun
实现方案
Google Cloud Speech-to-Text API默认生成的SRT是段落级的,但API已经返回了每个单词的时间戳(代码中已开启enable_word_time_offsets=True),因此可以通过解析单词级时间数据,自行生成细粒度SRT。
修改后的完整代码
from google.api_core.client_options import ClientOptions from google.cloud.speech_v2 import SpeechClient from google.cloud.speech_v2.types import cloud_speech import json from google.protobuf.json_format import MessageToDict MAX_AUDIO_LENGTH_SECS = 8 * 60 * 60 def format_srt_time(seconds): """将秒数转换为SRT格式的时间字符串:HH:MM:SS,mmm""" hours = int(seconds // 3600) minutes = int((seconds % 3600) // 60) secs = seconds % 60 milliseconds = int((secs - int(secs)) * 1000) return f"{hours:02d}:{minutes:02d}:{int(secs):02d},{milliseconds:03d}" def generate_fine_grained_srt(words, max_words_per_entry=2): """根据单词列表生成细粒度SRT,每个条目最多包含max_words_per_entry个单词""" srt_lines = [] entry_count = 1 i = 0 total_words = len(words) while i < total_words: # 确定当前条目的单词范围(1或2个) end_idx = min(i + max_words_per_entry, total_words) current_words = words[i:end_idx] # 获取条目的起始和结束时间 start_time = current_words[0]['startTime'] end_time = current_words[-1]['endTime'] # 生成SRT条目 srt_lines.append(str(entry_count)) srt_lines.append(f"{format_srt_time(start_time)} --> {format_srt_time(end_time)}") srt_lines.append(' '.join([word['word'] for word in current_words])) srt_lines.append('') # 空行分隔条目 entry_count += 1 i = end_idx return '\n'.join(srt_lines) def run_batch_recognize(): # 初始化客户端 client = SpeechClient( client_options=ClientOptions( api_endpoint="us-central1-speech.googleapis.com", ), ) # 音频文件的GCS地址 audio_gcs_uri = "<redacted>" config = cloud_speech.RecognitionConfig( explicit_decoding_config=cloud_speech.ExplicitDecodingConfig( encoding=cloud_speech.ExplicitDecodingConfig.AudioEncoding.LINEAR16, sample_rate_hertz=24000, audio_channel_count=1, ), features=cloud_speech.RecognitionFeatures( enable_word_confidence=True, enable_word_time_offsets=True, enable_automatic_punctuation=True, max_alternatives=1, # 只保留最优结果,减少处理量 ), model="short", language_codes=["en-US"], ) output_config = cloud_speech.RecognitionOutputConfig( inline_response_config=cloud_speech.InlineOutputConfig(), ) files = [cloud_speech.BatchRecognizeFileMetadata(uri=audio_gcs_uri)] request = cloud_speech.BatchRecognizeRequest( recognizer="<redacted>", config=config, files=files, recognition_output_config=output_config, ) operation = client.batch_recognize(request=request) print("等待识别完成...") response = operation.result(timeout=3 * MAX_AUDIO_LENGTH_SECS) # 转换为字典格式 response_dict = MessageToDict(response._pb) # 提取单词级数据 all_words = [] for result in response_dict["results"][audio_gcs_uri]["inlineResult"]["transcriptResults"][0]["alternatives"][0]["words"]: # 解析时间(API返回的是字符串,比如"1.234s",需要转换为秒数) start_time = float(result["startTime"].replace("s", "")) end_time = float(result["endTime"].replace("s", "")) all_words.append({ "word": result["word"], "startTime": start_time, "endTime": end_time }) # 生成细粒度SRT fine_grained_srt = generate_fine_grained_srt(all_words, max_words_per_entry=2) # 输出结果 print("细粒度SRT字幕:\n") print(fine_grained_srt) run_batch_recognize()
关键修改说明
- 时间格式转换函数:
format_srt_time将API返回的秒数转换为SRT标准的HH:MM:SS,mmm格式。 - 细粒度SRT生成函数:
generate_fine_grained_srt遍历单词列表,按1-2个单词为一组生成SRT条目,每组的起始时间取第一个单词的开始时间,结束时间取最后一个单词的结束时间。 - 结果解析逻辑:不再直接使用API生成的段落级SRT,而是提取
words字段中的单词级时间和文本数据,自行处理。 - 优化配置:将
max_alternatives改为1,只处理最优识别结果,减少不必要的数据处理。
效果示例
运行修改后的代码,会生成如下格式的细粒度SRT:
1 00:00:00,040 --> 00:00:00,540 The 2 00:00:00,540 --> 00:00:01,040 sun 3 00:00:01,040 --> 00:00:01,540 set
内容的提问来源于stack exchange,提问作者theprogrammer
相关产品推荐
相关产品推荐

