You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Cloud Speech-to-Text long_running_recognize结果无法迭代求助

解决Google Cloud Speech-to-Text长语音识别的字幕生成问题

调用long_running_recognize时,返回的是Operation异步操作对象,而非直接的识别响应。你需要先等待操作完成,再通过result()方法获取真正的LongRunningRecognizeResponse对象——这个对象的结构和短语音识别返回的RecognizeResponse完全兼容,同样包含可迭代的results属性。

完整实现代码

# 发起长语音识别异步请求
operation = client.long_running_recognize(config=config, audio=audio_uri)

# 等待操作完成(可根据音频长度调整timeout值,单位:秒)
response = operation.result(timeout=180)

subs_list = []
# 遍历结果的逻辑和短语音完全一致
for result in response.results:
    for alternative in result.alternatives:
         for word in alternative.words:
               # 处理start_time为空的情况
               start = word.start_time.total_seconds() if word.start_time else 0
               end = word.end_time.total_seconds()
               t = word.word
               subs_list.append( ((float(start), float(end)), t) )

print(subs_list)

关键说明

  1. 异步操作等待:long_running_recognize是异步接口,提交请求后不会立刻返回结果,必须通过operation.result()等待任务完成。如果音频时长较长,记得设置足够大的timeout参数,避免超时报错。
  2. 响应结构一致性:LongRunningRecognizeResponse和短语音的RecognizeResponse共享相同的结果结构,所以你可以直接复用短语音处理时的遍历逻辑,不需要修改循环结构。
  3. 错误原因:之前调用long_response.results报错,是因为long_response是Operation对象,它没有results属性,只有调用result()后得到的响应对象才有该属性。

内容的提问来源于stack exchange,提问作者Gablo Ficazzo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 02:35:24