借助Google API将语音识别转VTT字幕,获取时间戳问题求助
利用Google语音API获取带时间戳的识别结果并生成VTT字幕
核心修改点
要获取时间戳,必须在RecognitionConfig中开启词级时间戳;若音频时长超过1分钟,建议改用异步识别方法避免超时。
修改后的代码示例
var speech = SpeechClient.Create(); var config = new RecognitionConfig { AudioChannelCount = 2, Encoding = RecognitionConfig.Types.AudioEncoding.Flac, SampleRateHertz = 44100, LanguageCode = LanguageCodes.English.UnitedStates, // 开启词级时间戳,这是获取时间信息的关键配置 EnableWordTimeOffsets = true }; var audio = RecognitionAudio.FromFile(filepath); // 短音频用同步识别,长音频建议用下方注释的异步方法 var response = speech.Recognize(config, audio); // 生成VTT格式字幕内容 var vttContent = new List<string>(); vttContent.Add("WEBVTT"); vttContent.Add(""); int captionId = 1; foreach (var result in response.Results) { foreach (var alternative in result.Alternatives) { if (alternative.Words.Count == 0) continue; // 取当前识别片段的起始(第一个词的开始)和结束(最后一个词的结束)时间 var start = alternative.Words[0].StartTime; var end = alternative.Words[^1].EndTime; // 转换为VTT要求的时间格式:HH:MM:SS.mmm string startStr = $"{start.Hours:D2}:{start.Minutes:D2}:{start.Seconds:D2}.{start.Nanos / 1000000:D3}"; string endStr = $"{end.Hours:D2}:{end.Minutes:D2}:{end.Seconds:D2}.{end.Nanos / 1000000:D3}"; // 添加VTT条目 vttContent.Add($"{captionId}"); vttContent.Add($"{startStr} --> {endStr}"); vttContent.Add(alternative.Transcript); vttContent.Add(""); captionId++; } } // 将字幕写入文件 File.WriteAllLines("output.vtt", vttContent); Console.WriteLine("VTT字幕已生成");
补充说明
- 长音频处理:若音频时长超过1分钟,替换同步识别为异步方法避免超时,代码如下:
var operation = speech.LongRunningRecognize(config, audio); operation = operation.PollUntilCompleted(); var response = operation.Result;
- 时间格式转换:Google API返回的
Duration类型时间,需通过Nanos / 1000000计算出毫秒部分,匹配VTT格式要求。
内容的提问来源于stack exchange,提问作者Abdullah Mohamed
相关产品推荐
相关产品推荐

