You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

httpx异步HTTP2请求遇过载错误:如何实现降级重试不丢任务?

GitLab过载导致连接中断的解决方案

核心问题分析

当前使用httpx.AsyncClient异步请求GitLab API时,服务器过载会触发ConnectionTerminated错误(固定last_stream_id),导致任务直接返回None丢失测试报告数据;异步请求速率80req/sec在低负载时正常,但40+用户同时使用GitLab时,服务器无法承受瞬时并发量。


1. 请求降速:控制并发量

批量创建所有异步任务会导致瞬时请求量过高,用asyncio.Semaphore限制同时发起的请求数,避免压垮GitLab服务器:

修改FastAPI路由代码:

# 限制并发请求数,根据实际情况调整(比如设为20)
semaphore = asyncio.Semaphore(20)

async def bounded_testreport(project_id, pipeline_id, summary):
    async with semaphore:
        return await gitlab.testreport(project_id, pipeline_id, summary)

# 替换原任务创建逻辑
tasks = [bounded_testreport(v['project_id'], v['pipeline_id'], summary=True) for j in jobs]
for ta in asyncio.as_completed(tasks):
    testreports = await asyncio.gather(ta)
    for tr in testreports:
        <update job structure with testreports>

2. 重试机制:自动重试失败请求

使用tenacity库实现重试逻辑,针对ConnectionTerminated、5xx HTTP错误进行指数退避重试,避免任务直接丢失:

先安装依赖:pip install tenacity

修改gitlab.testreport函数:

from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type, retry_if_result
import httpx
from typing import Optional, Dict

# 全局会话初始创建
s = httpx.AsyncClient(http1=False, http2=True)

def is_none_result(result):
    return result is None

@retry(
    stop=stop_after_attempt(3),  # 最多重试3次
    wait=wait_exponential(multiplier=1, min=2, max=10),  # 指数退避:2s→4s→8s,最大间隔10s
    retry=(retry_if_exception_type(httpx.ConnectionTerminated) | 
           retry_if_exception_type(httpx.HTTPStatusError) |  # 捕获5xx等状态错误
           retry_if_result(is_none_result))  # 针对返回None的情况重试
)
async def testreport(id:int, pipid:int, summary:bool=False) -> Optional[Dict]:
    global s
    url = config.get_config()['gitlaburl'] + f"/{id}/pipelines/{pipid}/test_report"
    try:
        r = await s.get(url, headers={'PRIVATE-TOKEN': config.get_config()['gitlabtoken']})
        r.raise_for_status()
        # 成功后的处理逻辑
        data = r.json()
        return data if not summary else <summary处理逻辑>
    except httpx.ConnectionTerminated as exc:
        # 连接中断时重建会话,避免复用失效连接
        s = httpx.AsyncClient(http1=False, http2=True)
        log.error(f"GitLab连接中断,重建会话并重试: {exc.request.url} - {exc}")
        raise  # 抛出异常触发重试
    except httpx.HTTPError as exc:
        log.error(f"GitLab请求失败: {exc.request.url} - {exc}")
        return None

3. 全局会话重连:修复异常连接

当触发ConnectionTerminated时,当前httpx.AsyncClient的HTTP/2连接大概率处于异常状态,直接重建会话可以避免后续请求复用失效连接,这一步已经整合到上面的重试逻辑中。


4. 容错优化:避免任务彻底丢失

在FastAPI路由中收集失败的任务,进行二次重试,或者记录失败的项目/流水线ID,后续异步补全:

failed_tasks = []

tasks = [bounded_testreport(v['project_id'], v['pipeline_id'], summary=True) for j in jobs]
for idx, ta in enumerate(asyncio.as_completed(tasks)):
    try:
        testreports = await asyncio.gather(ta)
        for tr in testreports:
            if tr is None:
                failed_tasks.append(jobs[idx])
            else:
                <update job structure with testreports>
    except Exception as exc:
        failed_tasks.append(jobs[idx])
        log.error(f"任务执行失败: {exc}")

# 对失败任务进行二次重试
if failed_tasks:
    log.info(f"开始重试{len(failed_tasks)}个失败任务")
    retry_tasks = [bounded_testreport(v['project_id'], v['pipeline_id'], summary=True) for v in failed_tasks]
    for ta in asyncio.as_completed(retry_tasks):
        testreports = await asyncio.gather(ta)
        for tr in testreports:
            if tr is not None:
                <update job structure with testreports>

额外建议

  • 监控GitLab API速率限制:通过响应头X-RateLimit-Remaining和X-RateLimit-Reset获取当前速率限额,动态调整并发量,避免触发限流。
  • 缓存历史报告:将已获取的测试报告缓存到Redis或本地数据库,避免重复请求GitLab API,降低服务器负载。

内容的提问来源于stack exchange,提问作者MortenB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 17:42:48