使用aiohttp批量请求:复用ClientSession能否提升性能?
复用aiohttp ClientSession确实能大幅提速,你的问题出在用法和并发控制上
首先明确:复用ClientSession是提升批量请求速度的核心优化点,原代码每个请求都新建Session的做法完全是反模式——每次新建Session都会重新建立TCP连接、完成TLS握手,300万次请求的话,这部分开销会直接拖垮整体速度。
你的初步测试没看到效果,大概率是复用姿势不对,或者忽略了并发控制导致资源耗尽,掩盖了Session复用的收益。
正确的Session复用方式
把ClientSession作为客户端类的实例变量,初始化时创建一次,所有请求共用这个Session,用完再统一关闭:
import aiohttp import asyncio class MyClient: def __init__(self): # 初始化时创建唯一的ClientSession self.session = aiohttp.ClientSession() async def create(self, single_request_body: dict) -> bool: """请求成功返回True""" async with self.session.post( "https://my-server.org/endpoint", json=single_request_body # 如果是JSON接口,用json参数比data更高效 ) as response: return response.status == 201 async def close(self): # 客户端实例用完后关闭Session await self.session.close()
调用时必须控制并发数
直接用asyncio.gather一次性启动300万协程是致命错误:会瞬间耗尽系统文件句柄、内存,甚至触发服务器的限流机制,反而让速度更慢。必须用信号量(Semaphore)限制并发数:
async def main(): # 生成300万个不同的请求体(示例) all_request_bodies = [{"data": f"item_{i}"} for i in range(3_000_000)] my_client = MyClient() try: # 限制并发数,比如设为100(根据服务器承受能力调整) semaphore = asyncio.Semaphore(100) async def bounded_request(body): async with semaphore: return await my_client.create(body) # 批量创建任务并执行 tasks = [bounded_request(body) for body in all_request_bodies] results = await asyncio.gather(*tasks) # 可以统计成功/失败数量 success_count = sum(results) print(f"成功请求数:{success_count},失败请求数:{len(results)-success_count}") finally: await my_client.close() if __name__ == "__main__": asyncio.run(main())
额外优化建议
- 连接池配置:创建Session时可以自定义TCP连接池,优化连接复用效率:
connector = aiohttp.TCPConnector(limit=100, limit_per_host=100) self.session = aiohttp.ClientSession(connector=connector) - 请求体序列化:如果接口接收JSON,一定要用
json参数而不是data——data会把dict编码为form-data,序列化开销更大,也不符合多数API的预期格式。 - 分批处理:如果300万请求一次性处理内存压力大,可以分成若干批次(比如每批10000个),处理完一批再处理下一批,降低内存占用。
内容的提问来源于stack exchange,提问作者koks der drache
相关产品推荐
相关产品推荐

