You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI复用HttpClient响应延迟升高问题求助

问题描述

我有一个仅负责转发请求并返回响应的FastAPI端点。为避免频繁新建连接对目标服务器造成压力,采用共享HttpClient方案,但发现复用aiohttp或httpx的ClientSession时,响应时间从6-8ms飙升至30-50ms,且在每秒10-50次的高吞吐量场景下差异更明显,低并发时差异较小。无论通过Depends注入还是直接调用共享客户端,结果一致。对比Rust/Go中复用HttpClient会提升性能的情况,怀疑是Python的上下文切换导致?已尝试仅保留单个端点、将客户端单独存放等操作,使用Python 3.11.6,在Sanic框架中测试也得到类似结果,现有资料未提及该延迟问题,寻求原因排查及解决方法。

代码示例
httpx_client = httpx.AsyncClient()
session = requests.Session()
session.mount("http://", HTTPAdapter(pool_connections=1000, pool_maxsize=1000))
aiohttp_session = aiohttp.ClientSession()

def get_requests_session():
    yield session

async def get_aiohttp_session():
    yield aiohttp_session

app = FastAPI()
router = APIRouter()

# slow response time(40ms), reuses session
@router.get("/funtime")
async def search(
    query: str,
    limit: int,
    fa_response: Response,
    session: aiohttp.ClientSession = Depends(get_aiohttp_session),
):
    args = {"q": query, "limit": limit, "attributesToRetrieve": ["id", "title"]}
    url = config.base + "/indexes/pages/search"
    try:
        async with session.post(url, json=args) as response:
            if response.status != 200:
                raise HTTPException(status_code=response.status, detail="Error fetching data")
            return await response.json()
    except Exception as e:
        fa_response.status_code = status.HTTP_300_MULTIPLE_CHOICES
        return {"error": str(e), "url": url}

# slow response time (40ms), reuses session
@router.get("/funtime6")
async def search6(query: str, limit: int, fa_response: Response):
    args = {"q": query, "limit": limit, "attributesToRetrieve": ["id", "title"]}
    url = config.base + "/indexes/pages/search"

    try:
        response = await config.config_httpx_client.post(
            config.clerk_search_url + "/indexes/pages/search", json=args
        )
        if response.status_code != 200:
            raise HTTPException(status_code=response.status_code, detail="Error fetching data")
        hits = response.json().get("hits", [])
        return {"results": hits}

    except Exception as e:
        fa_response.status_code = status.HTTP_300_MULTIPLE_CHOICES
        return {"error": str(e), "url": url}

# fast response time (6ms), new session for each request
@router.get("/v3/test/funtime9")
def search9(query: str, limit: int, fa_response: Response):
    args = {"q": query, "limit": limit, "attributesToRetrieve": ["id", "title"]}
    url = config.base + "/indexes/pages/search"

    try:
        response = requests.post(url, json=args)
        if response.status_code != 200:
            raise HTTPException(status_code=response.status_code, detail="Error fetching data")
        hits = response.json().get("hits", [])
        return {"results": hits}
    except HTTPException as e:
        fa_response.status_code = status.HTTP_300_MULTIPLE_CHOICES
        return {"error": str(e), "url": url}

app.include_router(router)
原因排查与解决方法

核心原因分析

  1. 异步客户端连接池阻塞:aiohttp/httpx的ClientSession默认连接池大小较小(如aiohttp默认100),高并发下请求会排队等待可用连接,反而比每次新建连接更慢。而requests的同步请求用线程池处理,每个请求占独立线程,不会在连接池排队。
  2. 异步IO调度开销放大:Python的asyncio是单线程协程调度,高并发下协程切换开销会被放大,尤其是目标服务响应较快时,调度开销占比显著提升。同步requests的多线程模型在IO等待时会释放GIL,实际调度开销更低。
  3. 客户端配置未优化:共享ClientSession可能未针对高并发调整参数,比如未设置合理的连接超时、重试策略,或未开启TCP_NODELAY导致小数据包传输延迟。

具体解决措施

  • 调大连接池参数:
    初始化异步客户端时显式设置更大的连接池,避免请求排队:
    # aiohttp配置示例
    aiohttp_session = aiohttp.ClientSession(connector=aiohttp.TCPConnector(limit=1000, limit_per_host=1000))
    # httpx配置示例
    httpx_client = httpx.AsyncClient(limits=httpx.Limits(max_connections=1000, max_keepalive_connections=1000))
    
  • 启用TCP_NODELAY:
    关闭Nagle算法,减少小数据包传输延迟:
    connector = aiohttp.TCPConnector(limit=1000, tcp_nodelay=True)
    aiohttp_session = aiohttp.ClientSession(connector=connector)
    
  • 替换高效事件循环:
    用uvloop替代默认asyncio事件循环,它基于C实现,调度效率更高:
    import uvloop
    uvloop.install()
    
  • 优化同步客户端连接池:
    如果异步方案仍无改善,可优化requests的连接池,同时用多进程部署(如gunicorn)充分利用多核:
    session = requests.Session()
    adapter = HTTPAdapter(pool_connections=1000, pool_maxsize=1000, max_retries=3)
    session.mount("http://", adapter)
    session.mount("https://", adapter)
    
  • 排查目标服务连接限制:
    确认目标服务器是否对长连接有特殊限制(如超时过短、连接数上限),若目标服务频繁断开长连接,复用连接反而会增加重连开销。

内容的提问来源于stack exchange,提问作者LGXerxes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 17:25:53