You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

提升Python REST请求效率:关联数据请求性能问题问询

问题描述

之前通过SQL关联查询获取数据,10000条数据仅需几分之一秒,SQL语句如下:

select name,
       profession.expertise,
       resonsible.name,
       building.locker
from technicians
join profession on technicin.profession = profession.id
join responsible on techician.responsible = responsible.id
join building on technician.locker = buiding.id

迁移至云端后无法直接访问数据库,只能通过REST接口获取数据。但把上述逻辑转成Python REST请求后,仅50条数据就耗时20秒,10000条更是要数小时。原因是需要遍历所有结果,逐个请求关联数据,伪代码如下:

result_list = requests.get("https://sample.com/api/v1/technicians")
for technician in result_list:
    professn = requests.get("https://sample.com/api/v1/technicians/technician_ID/profession")
    resopnsible = requests.get("https://sample.com/api/v1/technicians/technician_ID/responsible")
    building = requests.get("https://sample.com/api/v1/technicians/technician_ID/building")

请问这种请求方式是否存在问题?

问题分析与解决方案

你的请求方式存在严重性能问题,核心是触发了经典的N+1查询陷阱:

  • 先获取N条技师数据,再为每条数据发起3次独立HTTP请求,总请求量达到1 + 3*N次。HTTP请求本身包含握手、传输等固定开销,当N=10000时,30001次请求的累积耗时会被无限放大。

针对这个问题,可按优先级尝试以下优化方案:

1. 要求API提供关联数据的批量/嵌入返回

这是最优解,直接从根源减少请求次数:

  • 询问API维护方,是否支持在技师列表接口中通过参数(比如?include=profession,responsible,building)直接返回关联数据,像你原来的SQL JOIN那样一次性拿到所有需要的信息,把请求次数降到1次。
  • 如果支持批量查询关联资源(比如/api/v1/professions/batch允许传入多个ID批量获取),可以先收集所有技师的关联ID,再分批次发起批量请求,将请求次数压缩到1 + 3次(或少量批次)。

2. 用异步请求并发处理关联查询

如果API不支持批量或嵌入,可改用异步请求库(如aiohttp)替代同步的requests,并发发起关联数据请求,消除串行等待的时间损耗:

import aiohttp
import asyncio

async def fetch_related(session, tech_id):
    tasks = [
        session.get(f"https://sample.com/api/v1/technicians/{tech_id}/profession"),
        session.get(f"https://sample.com/api/v1/technicians/{tech_id}/responsible"),
        session.get(f"https://sample.com/api/v1/technicians/{tech_id}/building")
    ]
    return await asyncio.gather(*tasks)

async def main():
    async with aiohttp.ClientSession() as session:
        tech_resp = await session.get("https://sample.com/api/v1/technicians")
        result_list = await tech_resp.json()
        
        tasks = [fetch_related(session, tech['id']) for tech in result_list]
        await asyncio.gather(*tasks)

asyncio.run(main())

这种方式仍会产生3N次请求,但并发执行能大幅缩短总耗时。

3. 分页获取+并发结合

如果技师列表支持分页(比如?page=1&size=100),可以分页拉取技师数据,同时对每一页的技师并发请求关联数据,既提升效率,也避免一次性发起大量请求触发API限流。


内容的提问来源于stack exchange,提问作者hewi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 01:03:24