You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google Speech-to-Text API处理长音频时重试过多失败求助

解决Google Speech-to-Text处理长音频重试过多失败的问题

我之前处理长音频转写时也碰到过一模一样的重试失败问题,结合官方文档和实际调试经验,给你几个可行的解决方向:

1. 自定义客户端重试策略,放宽重试限制

默认的Google Cloud客户端重试策略对长时任务不够友好,你可以通过自定义重试参数,增加重试次数、延长重试间隔,适配长音频的处理耗时:

from google.api_core.retry import Retry, exponential_backoff
from google.api_core.exceptions import GoogleAPICallError, DeadlineExceeded, ServiceUnavailable

# 针对长时转写任务定义更宽松的重试策略
custom_retry = Retry(
    # 指定需要重试的异常类型
    predicate=lambda exc: isinstance(exc, (GoogleAPICallError, DeadlineExceeded, ServiceUnavailable)),
    # 指数退避策略:初始间隔10秒,每次翻倍,最大间隔60秒
    backoff=exponential_backoff(initial=10, multiplier=2, maximum=60),
    maximum_attempts=15  # 把默认的5次重试提高到15次
)

# 调用long_running_recognize时传入自定义重试策略
operation = client.long_running_recognize(config, audio, retry=custom_retry)

2. 优化任务等待逻辑,避免冲突轮询

你的代码里同时用了add_done_callback、自定义percentile轮询和time.sleep(30),可能会导致重复请求加重API负载。建议简化等待逻辑:

def callback(operation_future):
    try:
        result = operation_future.result()
        print("✅ 音频转写任务完成")
        # 这里可以添加结果处理的业务逻辑
    except Exception as e:
        print(f"❌ 任务失败: {str(e)}")

# 添加回调后直接等待结果,无需额外sleep和自定义轮询
operation.add_done_callback(callback)
print("等待转写任务完成...")
response = operation.result(timeout=None)  # timeout=None表示无限等待直到任务结束

3. 分割长音频为短片段处理(终极兜底方案)

如果调整重试策略后还是失败,最稳妥的方式是把超过60分钟的音频分割成多个30-60分钟的片段,分别处理后合并结果:

用FFmpeg分割OGG OPUS音频

# 分割为每个30分钟的片段(1800秒),保持编码不变
ffmpeg -i your_audio.ogg -f segment -segment_time 1800 -c copy audio_segment_%03d.ogg

代码中批量处理并合并结果

处理每个片段时,记得给每个词的时间偏移量加上片段的起始时间,保证最终字幕时间轴正确:

import os
from datetime import timedelta

# 遍历所有分割后的音频片段
segment_files = sorted([f for f in os.listdir('.') if f.startswith('audio_segment_')])
full_transcript = []

for idx, segment in enumerate(segment_files):
    segment_uri = f"gs://your-bucket/{segment}"
    audio = types.RecognitionAudio(uri=segment_uri)
    operation = client.long_running_recognize(config, audio)
    segment_result = operation.result(timeout=None)
    
    # 计算当前片段的起始时间(30分钟 * 片段序号)
    start_offset = timedelta(minutes=30 * idx)
    
    # 调整每个词的时间偏移
    for result in segment_result.results:
        for alternative in result.alternatives:
            for word_info in alternative.words:
                word_info.start_time += start_offset
                word_info.end_time += start_offset
        full_transcript.append(result)

# full_transcript就是合并后的完整转写结果

4. 检查音频文件完整性

最后别忘了确认GCS上的音频文件没有损坏,OGG OPUS编码是否正确,采样率确实是48000Hz——文件损坏也可能导致API处理超时,触发重试过多的错误。

内容的提问来源于stack exchange,提问作者DLI42

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:38:42