使用Google Speech-to-Text API处理长音频时重试过多失败求助
解决Google Speech-to-Text处理长音频重试过多失败的问题
我之前处理长音频转写时也碰到过一模一样的重试失败问题,结合官方文档和实际调试经验,给你几个可行的解决方向:
1. 自定义客户端重试策略,放宽重试限制
默认的Google Cloud客户端重试策略对长时任务不够友好,你可以通过自定义重试参数,增加重试次数、延长重试间隔,适配长音频的处理耗时:
from google.api_core.retry import Retry, exponential_backoff from google.api_core.exceptions import GoogleAPICallError, DeadlineExceeded, ServiceUnavailable # 针对长时转写任务定义更宽松的重试策略 custom_retry = Retry( # 指定需要重试的异常类型 predicate=lambda exc: isinstance(exc, (GoogleAPICallError, DeadlineExceeded, ServiceUnavailable)), # 指数退避策略:初始间隔10秒,每次翻倍,最大间隔60秒 backoff=exponential_backoff(initial=10, multiplier=2, maximum=60), maximum_attempts=15 # 把默认的5次重试提高到15次 ) # 调用long_running_recognize时传入自定义重试策略 operation = client.long_running_recognize(config, audio, retry=custom_retry)
2. 优化任务等待逻辑,避免冲突轮询
你的代码里同时用了add_done_callback、自定义percentile轮询和time.sleep(30),可能会导致重复请求加重API负载。建议简化等待逻辑:
def callback(operation_future): try: result = operation_future.result() print("✅ 音频转写任务完成") # 这里可以添加结果处理的业务逻辑 except Exception as e: print(f"❌ 任务失败: {str(e)}") # 添加回调后直接等待结果,无需额外sleep和自定义轮询 operation.add_done_callback(callback) print("等待转写任务完成...") response = operation.result(timeout=None) # timeout=None表示无限等待直到任务结束
3. 分割长音频为短片段处理(终极兜底方案)
如果调整重试策略后还是失败,最稳妥的方式是把超过60分钟的音频分割成多个30-60分钟的片段,分别处理后合并结果:
用FFmpeg分割OGG OPUS音频
# 分割为每个30分钟的片段(1800秒),保持编码不变 ffmpeg -i your_audio.ogg -f segment -segment_time 1800 -c copy audio_segment_%03d.ogg
代码中批量处理并合并结果
处理每个片段时,记得给每个词的时间偏移量加上片段的起始时间,保证最终字幕时间轴正确:
import os from datetime import timedelta # 遍历所有分割后的音频片段 segment_files = sorted([f for f in os.listdir('.') if f.startswith('audio_segment_')]) full_transcript = [] for idx, segment in enumerate(segment_files): segment_uri = f"gs://your-bucket/{segment}" audio = types.RecognitionAudio(uri=segment_uri) operation = client.long_running_recognize(config, audio) segment_result = operation.result(timeout=None) # 计算当前片段的起始时间(30分钟 * 片段序号) start_offset = timedelta(minutes=30 * idx) # 调整每个词的时间偏移 for result in segment_result.results: for alternative in result.alternatives: for word_info in alternative.words: word_info.start_time += start_offset word_info.end_time += start_offset full_transcript.append(result) # full_transcript就是合并后的完整转写结果
4. 检查音频文件完整性
最后别忘了确认GCS上的音频文件没有损坏,OGG OPUS编码是否正确,采样率确实是48000Hz——文件损坏也可能导致API处理超时,触发重试过多的错误。
内容的提问来源于stack exchange,提问作者DLI42
相关产品推荐
相关产品推荐

