You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法访问Google Speech识别结果:并发长时识别代码报错求助

解决长时语音识别任务的AttributeError问题

你碰到的这个问题其实很常见——你存到output字典里的是Google Cloud返回的Operation对象,而不是实际的语音识别结果!当job.done()返回True时,只是说明任务完成了,但要拿到转录内容,你得调用job.result()方法来提取真正的识别响应数据。

错误原因拆解

Google Cloud Speech-to-Text的long_running_recognize方法返回的是Operation对象,它负责跟踪任务进度、状态,以及最终存储结果。这个对象本身没有alternatives属性——alternatives是嵌套在LongRunningRecognizeResponse里的,而这个响应需要通过Operation.result()来获取。

修正后的代码

下面是调整后的完整代码,我标注了关键的修改点:

import time
from google.cloud import speech
from google.cloud.speech import enums, types

client = speech.SpeechClient()
config = types.RecognitionConfig(
    encoding=enums.RecognitionConfig.AudioEncoding.FLAC,
    language_code='en-US')
# 调整audio结构,确保每个键对应一个RecognitionAudio对象
audio = {
    "Brooklyn": types.RecognitionAudio(uri='gs://cloud-samples-tests/speech/brooklyn.flac')
}
jobs = {}
output = {}

for name, audio_obj in audio.items():
    jobs[name] = client.long_running_recognize(config, audio_obj)

while len(jobs) > 0:
    time.sleep(5)
    # 遍历字典副本避免修改时触发RuntimeError
    for name, job in list(jobs.items()):
        if not job.done():
            print(f'{name} progress: {job.metadata.progress_percent}%')
        else:
            print(f'{name} is done!')
            try:
                # 关键修改:调用result()获取实际识别响应
                response = job.result()
                output[name] = response
            except Exception as e:
                print(f'{name} failed with error: {e}')
            jobs.pop(name)

# 正确遍历识别结果的嵌套结构
for name, response in output.items():
    print(f'\n--- {name} Transcript ---')
    # 每个响应包含多个识别片段结果
    for result in response.results:
        top_alternative = result.alternatives[0]
        print(f'Transcript: {top_alternative.transcript}')
        print(f'Confidence: {top_alternative.confidence}')

额外注意点

  1. 字典遍历的坑:原代码直接遍历jobs.items()并执行pop会引发RuntimeError,改成遍历list(jobs.items())就可以避开这个问题——因为我们遍历的是字典的副本,修改原字典不会影响遍历过程。
  2. 异常防护:调用job.result()可能会抛出API错误(比如音频无效、权限不足),加上try-except可以让程序更健壮,还能捕获具体的错误信息。
  3. 结果结构细节:LongRunningRecognizeResponse的results是一个列表,每个元素对应音频中的一段连续识别内容,所以需要循环遍历这些结果才能获取完整的转录内容。

内容的提问来源于stack exchange,提问作者Ben Etherington

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:02:31