为何我的asyncio HTTP请求I/O等待未按预期重叠?
问题描述
尝试对同一API的不同端点发起多个HTTP GET请求,预期通过asyncio + aiohttp实现I/O等待重叠,使总耗时接近最慢单个请求的耗时,但实际总耗时约等于各请求耗时之和,请求呈现串行执行的状态。使用Python 3.12.0、aiohttp==3.9.3版本,代码中无明显遗漏await语句。
代码示例:
import asyncio import aiohttp import time async def fetch_data(session, url): """ Fetches data from a given URL using aiohttp. Includes print statements to track apparent start/finish times. """ start_req_time = time.perf_counter() print(f"[{time.time():.2f}] Starting request for: {url}") try: async with session.get(url) as response: response.raise_for_status() # Raise an exception for bad status codes data = await response.json() # Or .text(), depending on expected response end_req_time = time.perf_counter() print(f"[{time.time():.2f}] Finished request for: {url} in {end_req_time - start_req_time:.4f}s") return {"url": url, "status": response.status, "data_length": len(str(data))} except aiohttp.ClientError as e: print(f"[{time.time():.2f}] Error fetching {url}: {e}") return {"url": url, "error": str(e)} async def main(): # Using a list of public JSON placeholder URLs for demonstration # In my real app, these are different endpoints of my own API urls = [ "https://jsonplaceholder.typicode.com/todos/1", "https://jsonplaceholder.typicode.com/todos/2", "https://jsonplaceholder.typicode.com/posts/1", "https://jsonplaceholder.typicode.com/users/1", "https://jsonplaceholder.typicode.com/comments/1", ] total_start_time = time.perf_counter() async with aiohttp.ClientSession() as session: tasks = [] for url in urls: tasks.append(fetch_data(session, url)) # This is where I expect the I/O waits of tasks to overlap results = await asyncio.gather(*tasks) total_end_time = time.perf_counter() print(f"\n[{time.time():.2f}] All tasks completed in {total_end_time - total_start_time:.4f} seconds.") print("\n--- Results ---") for result in results: print(result) if __name__ == "__main__": asyncio.run(main())
运行输出模式:
[1678886400.10] Starting request for: https://jsonplaceholder.typicode.com/todos/1 [1678886400.25] Finished request for: https://jsonplaceholder.typicode.com/todos/1 in 0.1500s [1678886400.25] Starting request for: https://jsonplaceholder.typicode.com/todos/2 [1678886400.40] Finished request for: https://jsonplaceholder.typicode.com/todos/2 in 0.1500s [1678886400.40] Starting request for: https://jsonplaceholder.typicode.com/posts/1 [1678886400.55] Finished request for: https://jsonplaceholder.typicode.com/posts/1 in 0.1500s ... and so on ... [1678886400.80] All tasks completed in 0.7000 seconds.
核心问题:请求串行执行,总耗时接近各请求耗时之和,未实现预期的并发I/O重叠。
原因分析与解决方案
TCP连接池并发限制:aiohttp的
TCPConnector默认对同一主机的并发连接数(limit_per_host)设置为2,这意味着同一域名下最多同时处理2个请求,剩余请求会排队等待连接释放,导致视觉上的串行执行。解决办法:创建
ClientSession时显式指定TCPConnector的limit_per_host参数,提高同一主机的并发连接数:async with aiohttp.ClientSession(connector=aiohttp.TCPConnector(limit_per_host=10)) as session:可根据实际请求数量调整数值,确保所有请求能同时建立连接。
DNS解析阻塞(次要可能):如果系统默认的DNS解析是同步实现,可能会在建立连接前阻塞事件循环。可显式使用异步DNS解析器优化:
resolver = aiohttp.AsyncResolver() connector = aiohttp.TCPConnector(resolver=resolver, limit_per_host=10) async with aiohttp.ClientSession(connector=connector) as session:服务器端限流:如果目标API服务器本身限制了同一IP的并发请求数,即使客户端放开连接限制,服务器也会串行处理请求。这种情况需要确认API的限流规则,或调整请求速率适配规则。
修改后,所有请求会同时发起,总耗时将接近单个请求的最长耗时(如0.2秒左右)。
内容的提问来源于stack exchange,提问作者drawfine
相关产品推荐
相关产品推荐

