aiohttp+asyncio网页解析初始正常,运行后无数据返回求助
异步批量爬取域名主页无响应的修复方案
你的脚本初期运行正常,后期突然无数据返回,核心原因是一次性创建所有请求任务导致并发量过高:aiohttp的ClientSession默认连接池资源有限,瞬间大量请求会耗尽连接池,导致后续请求无法发起;同时高并发也容易触发目标网站反爬机制或网络限流,造成请求“静默失败”。
此外还有几个潜在问题:
- 全局字典
data在高并发场景下虽无线程安全问题,但异步任务同时写入可能引发数据覆盖或异常; - 仅使用
http://前缀,部分域名会强制跳转HTTPS,可能引发超时或错误; - 全局120秒超时设置过于宽泛,单个慢请求会阻塞整体流程。
解决方案
- 限制并发数:用
asyncio.Semaphore控制同时运行的任务数量,避免请求过载; - 避免全局共享状态:让任务返回结果,统一收集数据,替代全局字典;
- 支持HTTPS并优化请求逻辑:优先尝试HTTPS,适配主流网站配置;
- 细粒度超时控制:给单个请求设置独立超时,避免慢请求拖垮整体;
- 完善错误日志:方便快速定位请求失败原因。
修改后的代码
import asyncio import time import aiohttp import pandas as pd import json from bs4 import BeautifulSoup async def get_data(session, url, semaphore): async with semaphore: try: # 单个请求设置30秒超时,避免慢请求阻塞 async with session.get(url, timeout=aiohttp.ClientTimeout(total=30)) as resp: html = await resp.text() rawhtml = BeautifulSoup(html, 'lxml') title = rawhtml.title.string if rawhtml.title else 'None' return { 'url': url, 'resp': resp.status, 'title': title, 'html': str(rawhtml) } except Exception as e: print(f"请求失败: {url} | 错误: {str(e)}") return { 'url': url, 'resp': str(e), 'title': 'None', 'html': 'None' } async def main(): df = pd.read_excel('work_file.xlsx') domains_count = df.shape[0] print(f"总域名数: {domains_count}") # 限制并发数,根据网络情况调整,建议10-20 semaphore = asyncio.Semaphore(15) tasks = [] async with aiohttp.ClientSession() as session: for domain in df.domain: # 优先尝试HTTPS,适配主流网站配置 url = f'https://{domain.lower()}' task = asyncio.create_task(get_data(session, url, semaphore)) tasks.append(task) # 等待所有任务完成并收集结果 results = await asyncio.gather(*tasks) # 转为原脚本的字典格式(可选) data = {item['url']: item for item in results} # 写入JSON文件,指定编码和缩进提升可读性 with open('parsed_data.json', 'w', encoding='utf-8') as file: json.dump(data, file, ensure_ascii=False, indent=2) if __name__ == '__main__': start_time = time.time() asyncio.run(main()) end_time = time.time() - start_time print(f"总耗时: {end_time:.2f}秒")
关键修改说明
- 用
Semaphore限制并发数,避免瞬间请求过载导致连接池耗尽; get_data函数返回结果,在main中统一收集,避免全局共享状态的潜在问题;- 单个请求设置独立超时,不会因某个慢请求阻塞所有任务;
- 默认使用HTTPS,适配当前绝大多数网站的配置;
- 优化title提取逻辑,直接获取标题文本而非标签字符串;
- 写入JSON时指定UTF-8编码和缩进,提升文件可读性。
内容的提问来源于stack exchange,提问作者Artem Biryukov
相关产品推荐
相关产品推荐

