使用httpx异步请求大Zip文件为何未比同步请求提速?
异步httpx下载大文件耗时与同步请求相近的原因分析
问题场景
使用httpx发起25次异步请求下载约5MB的Zip文件时,发现耗时与同步requests请求基本一致,两种方式平均耗时相近,未体现出异步请求的优势。测试代码如下:
import asyncio import time import httpx import requests async def main(): n = 25 url = f'https://www2.census.gov/geo/tiger/TIGER2020/PLACE/tl_2020_01_place.zip' start_time = time.perf_counter() async with httpx.AsyncClient() as client: tasks = [client.get(url) for _ in range(n)] responses = await asyncio.gather(*tasks) end_time = time.perf_counter() print('httpx time:', end_time-start_time) start_time = time.perf_counter() session = requests.Session() responses = [session.get(url=url) for _ in range(n)] end_time = time.perf_counter() print('sync time', end_time-start_time) loop = asyncio.get_event_loop() loop.run_until_complete(main())
核心原因分析
1. 默认响应读取方式阻塞异步优势
httpx异步客户端的get()方法默认会一次性将完整响应内容加载到内存,这意味着每个异步任务都需要等待整个文件下载完成才会结束。异步编程的核心优势是在IO等待阶段(如等待服务器返回数据)切换执行其他任务,但这种全量读取的方式让异步任务失去了切换的机会,和同步请求的执行逻辑本质趋同——都是等待每个请求完整完成,因此耗时差距不大。
2. 服务器并发连接限制
目标服务器大概率对单个客户端的并发连接数做了限制(比如同一时间仅允许2-3个连接)。即使你的异步代码发起了25个请求,httpx默认连接池上限为10,再加上服务器的限制,大部分请求会处于排队等待连接的状态,实际并发数与同步请求的串行/少量并行效果接近,总耗时自然没有明显优势。
3. 同步代码复用连接抵消差距
同步代码使用了requests.Session(),它会复用TCP连接,减少了重复建立连接的耗时。而异步AsyncClient同样会复用连接,但服务器的并发限制抵消了这种复用带来的性能提升,导致两者耗时进一步接近。
优化方案
1. 改用流式下载释放异步优势
通过流式分块读取响应,让异步任务在等待数据返回时可以切换到其他任务执行,真正发挥异步IO的优势。修改后的代码示例:
import asyncio import time import httpx import requests async def async_download(client, url): async with client.get(url, stream=True) as response: content = b"" async for chunk in response.aiter_bytes(): content += chunk return content async def main(): n = 25 url = 'https://www2.census.gov/geo/tiger/TIGER2020/PLACE/tl_2020_01_place.zip' start_time = time.perf_counter() async with httpx.AsyncClient() as client: tasks = [async_download(client, url) for _ in range(n)] responses = await asyncio.gather(*tasks) end_time = time.perf_counter() print('httpx async streaming time:', end_time-start_time) start_time = time.perf_counter() session = requests.Session() responses = [] for _ in range(n): with session.get(url, stream=True) as response: content = b"" for chunk in response.iter_content(chunk_size=8192): content += chunk responses.append(content) end_time = time.perf_counter() print('sync streaming time', end_time-start_time) loop = asyncio.get_event_loop() loop.run_until_complete(main())
2. 调整连接池大小(谨慎使用)
可以在创建AsyncClient时增大连接池上限,突破默认限制,但需注意不要超过服务器允许的并发数,避免被封禁:
async with httpx.AsyncClient(limits=httpx.Limits(max_connections=20)) as client: # 执行异步下载任务
内容的提问来源于stack exchange,提问作者Ethan Singer
相关产品推荐
相关产品推荐

