Python aiohttp批量请求超5000条触发asyncio.TimeoutError如何解决
核心原因
- 你设置的
ClientTimeout(total=600)是整个ClientSession从创建到关闭的全局总超时,按照你限流器每秒2个请求的配置,5000条数据需要至少2500秒才能跑完,已经远超过600秒阈值,数据量超过5000后超时必然会触发。 - 你没有给单个请求设置独立超时,也没有做异常捕获,只要某一个请求被服务器挂起无响应,就会触发连锁超时,同时
asyncio.gather默认只要有一个任务抛出异常就会直接终止整个程序。 - 一次性创建了所有异步任务,哪怕有限流器控制实际请求发送速率,数千个待执行任务也会占用事件循环资源,同时部分任务等待调度时间过长也可能触发超时。
- 长时间高频请求同一站点,可能触发站点WAF防护,服务器会主动挂起你的连接不返回响应,导致超时。
解决方案
1. 调整超时配置
关闭ClientSession的全局超时,给单个请求设置独立超时,避免整体运行时间过长触发全局超时:
from aiohttp import ClientTimeout # 关闭session全局超时 session_timeout = ClientTimeout(total=0) # 单个请求超时时间设为30秒,可根据实际情况调整 per_req_timeout = ClientTimeout(total=30)
2. 增加异常捕获和状态校验
在单条数据采集逻辑中增加异常捕获和响应状态校验,避免单个请求失败导致整个程序终止:
async def fetch_one(client, i): url = f'https://example.com/product/{i}' try: async with await client.get(url, timeout=per_req_timeout) as resp: # 校验响应状态码,非200直接跳过 if resp.status != 200: print(f"请求{url}失败,状态码:{resp.status}") return resp_text = await resp.text() bs0bj = BeautifulSoup(resp_text, 'lxml') item_node = bs0bj.find('div', {'class': 'item_'}) if not item_node: print(f"页面{url}未匹配到目标元素") return code_EAN = re.sub('^\s+|\n|\r|\s+$', '', item_node.get_text()) code_VEND_list.append(code_EAN) except Exception as e: print(f"处理{url}出错:{str(e)}") # 可在此处增加重试逻辑,请求失败后重试2-3次再放弃
3. 分批次处理任务
不要一次性创建所有异步任务,改为分批次处理,降低事件循环压力,同时可以在批次之间增加休眠时间降低服务器压力:
async def main(): vendor_seorce = [] workbook = xlrd.open_workbook('Source.xlsx') worksheet = workbook.sheet_by_index(0) for vend_code in worksheet.col_values(0): vendor_seorce.append(vend_code) async with aiohttp.ClientSession(timeout=session_timeout) as client: client = RateLimiter(client) # 每批次处理200条,可根据实际情况调整 batch_size = 200 for batch_idx in range(0, len(vendor_seorce), batch_size): current_batch = vendor_seorce[batch_idx:batch_idx+batch_size] tasks = [asyncio.ensure_future(fetch_one(client, id_cat)) for id_cat in current_batch] # 加return_exceptions避免单个任务异常终止整个批次 await asyncio.gather(*tasks, return_exceptions=True) # 批次之间休眠2秒,降低服务器压力 await asyncio.sleep(2) finish_time = time.time() - start_time print('Done!') print(f'Time: {finish_time}') file_writher()
4. 可选优化
如果还是出现超时,可以适当降低限流器的请求速率,比如将RATE调整为1(每秒1次请求),避免触发站点的频率限制。
内容的提问来源于stack exchange,提问作者Klikovskiy
相关产品推荐
相关产品推荐

