httpx异步HTTP2请求遇过载错误:如何实现降级重试不丢任务?
GitLab过载导致连接中断的解决方案
核心问题分析
当前使用httpx.AsyncClient异步请求GitLab API时,服务器过载会触发ConnectionTerminated错误(固定last_stream_id),导致任务直接返回None丢失测试报告数据;异步请求速率80req/sec在低负载时正常,但40+用户同时使用GitLab时,服务器无法承受瞬时并发量。
1. 请求降速:控制并发量
批量创建所有异步任务会导致瞬时请求量过高,用asyncio.Semaphore限制同时发起的请求数,避免压垮GitLab服务器:
修改FastAPI路由代码:
# 限制并发请求数,根据实际情况调整(比如设为20) semaphore = asyncio.Semaphore(20) async def bounded_testreport(project_id, pipeline_id, summary): async with semaphore: return await gitlab.testreport(project_id, pipeline_id, summary) # 替换原任务创建逻辑 tasks = [bounded_testreport(v['project_id'], v['pipeline_id'], summary=True) for j in jobs] for ta in asyncio.as_completed(tasks): testreports = await asyncio.gather(ta) for tr in testreports: <update job structure with testreports>
2. 重试机制:自动重试失败请求
使用tenacity库实现重试逻辑,针对ConnectionTerminated、5xx HTTP错误进行指数退避重试,避免任务直接丢失:
先安装依赖:pip install tenacity
修改gitlab.testreport函数:
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type, retry_if_result import httpx from typing import Optional, Dict # 全局会话初始创建 s = httpx.AsyncClient(http1=False, http2=True) def is_none_result(result): return result is None @retry( stop=stop_after_attempt(3), # 最多重试3次 wait=wait_exponential(multiplier=1, min=2, max=10), # 指数退避:2s→4s→8s,最大间隔10s retry=(retry_if_exception_type(httpx.ConnectionTerminated) | retry_if_exception_type(httpx.HTTPStatusError) | # 捕获5xx等状态错误 retry_if_result(is_none_result)) # 针对返回None的情况重试 ) async def testreport(id:int, pipid:int, summary:bool=False) -> Optional[Dict]: global s url = config.get_config()['gitlaburl'] + f"/{id}/pipelines/{pipid}/test_report" try: r = await s.get(url, headers={'PRIVATE-TOKEN': config.get_config()['gitlabtoken']}) r.raise_for_status() # 成功后的处理逻辑 data = r.json() return data if not summary else <summary处理逻辑> except httpx.ConnectionTerminated as exc: # 连接中断时重建会话,避免复用失效连接 s = httpx.AsyncClient(http1=False, http2=True) log.error(f"GitLab连接中断,重建会话并重试: {exc.request.url} - {exc}") raise # 抛出异常触发重试 except httpx.HTTPError as exc: log.error(f"GitLab请求失败: {exc.request.url} - {exc}") return None
3. 全局会话重连:修复异常连接
当触发ConnectionTerminated时,当前httpx.AsyncClient的HTTP/2连接大概率处于异常状态,直接重建会话可以避免后续请求复用失效连接,这一步已经整合到上面的重试逻辑中。
4. 容错优化:避免任务彻底丢失
在FastAPI路由中收集失败的任务,进行二次重试,或者记录失败的项目/流水线ID,后续异步补全:
failed_tasks = [] tasks = [bounded_testreport(v['project_id'], v['pipeline_id'], summary=True) for j in jobs] for idx, ta in enumerate(asyncio.as_completed(tasks)): try: testreports = await asyncio.gather(ta) for tr in testreports: if tr is None: failed_tasks.append(jobs[idx]) else: <update job structure with testreports> except Exception as exc: failed_tasks.append(jobs[idx]) log.error(f"任务执行失败: {exc}") # 对失败任务进行二次重试 if failed_tasks: log.info(f"开始重试{len(failed_tasks)}个失败任务") retry_tasks = [bounded_testreport(v['project_id'], v['pipeline_id'], summary=True) for v in failed_tasks] for ta in asyncio.as_completed(retry_tasks): testreports = await asyncio.gather(ta) for tr in testreports: if tr is not None: <update job structure with testreports>
额外建议
- 监控GitLab API速率限制:通过响应头
X-RateLimit-Remaining和X-RateLimit-Reset获取当前速率限额,动态调整并发量,避免触发限流。 - 缓存历史报告:将已获取的测试报告缓存到Redis或本地数据库,避免重复请求GitLab API,降低服务器负载。
内容的提问来源于stack exchange,提问作者MortenB
相关产品推荐
相关产品推荐

