使用aiohttp异步爬虫遇AttributeError: NoneType无connect属性问题求助
我正在编写针对freelance.habr.com的异步解析器,遍历特定查询的所有任务,使用aiohttp发起请求。执行过程中仅部分任务出现错误,且每次后续启动时错误数量都会增加。另外,错误信息指向系统Python路径而非虚拟环境路径,对此感到疑惑。
代码实现
import asyncio import aiohttp from bs4 import BeautifulSoup from fake_useragent import UserAgent async def get_data_from_habr() -> None: """Main function, adding info from all pages to database.""" p = 1 async with aiohttp.ClientSession() as session: while True: url = f"https://freelance.habr.com/tasks?categories=development_bots&page={p}" headers = {"user-agent": UserAgent().random} async with session.get(url=url, headers=headers) as res: src = await res.text() soup = BeautifulSoup(src, "lxml") if soup.find(class_="empty-block__title"): break orders = soup.find_all(class_="task__title") async_tasks = [] for order in orders: order_url = "https://freelance.habr.com" + order.find("a")["href"] async_tasks.append( asyncio.create_task(get_data_from_habr_order_page(order_url, session)) ) p += 1 await asyncio.gather(*async_tasks) async def get_data_from_habr_order_page(order_url: str, session: aiohttp.ClientSession) -> None: """Functions for getting info from one page""" headers = {"user-agent": UserAgent().random} async with session.get(url=order_url, headers=headers, timeout=1000) as res: src = await res.text()
错误信息
Traceback (most recent call last): File "\OrderScraper\habr_scraper.py", line 92, in <module> asyncio.run(get_data_from_habr()) File "\AppData\Local\Programs\Python\Python312\Lib\asyncio\runners.py", line 194, in run return runner.run(main) ^^^^^^^^^^^^^^^^ File "\AppData\Local\Programs\Python\Python312\Lib\asyncio\runners.py", line 118, in run return self._loop.run_until_complete(task) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "\AppData\Local\Programs\Python\Python312\Lib\asyncio\base_events.py", line 687, in run_until_complete return future.result() ^^^^^^^^^^^^^^^ File "\OrderScraper\habr_scraper.py", line 39, in get_data_from_habr await asyncio.gather(*async_tasks) File "\OrderScraper\habr_scraper.py", line 46, in get_data_from_habr_order_page async with session.get(url=order_url, headers=headers, timeout=1000) as res: File "\OrderScraper\order_scraper\Lib\site-packages\aiohttp\client.py", line 1353, in __aenter__ self._resp = await self._coro ^^^^^^^^^^^^^^^^ File "\OrderScraper\order_scraper\Lib\site-packages\aiohttp\client.py", line 657, in _request conn = await self._connector.connect( ^^^^^^^^^^^^^^^^^^^^^^^ AttributeError: 'NoneType' object has no attribute 'connect'
1. 核心错误:ClientSession提前关闭
代码中await asyncio.gather(*async_tasks)被放在了async with aiohttp.ClientSession() as session:的上下文管理器外部——当while循环因遇到空页面break后,session会被自动关闭(上下文管理器退出时销毁资源),此时再去await任务,任务中调用session.get()必然失败:session内部的connector已经被置为None,所以触发AttributeError: 'NoneType' object has no attribute 'connect'。
同时,async_tasks在while循环内每次都会被重置为空列表,最后gather的只有最后一页的任务;而前面页面创建的任务没有被await,会在后台运行,当session关闭后这些任务也会因session失效报错,这就是错误数量随每次启动增加的原因(未被正确await的后台任务积累)。
2. 系统Python路径显示的原因
错误栈中出现系统Python路径是正常现象:asyncio是Python标准库,不会被安装到虚拟环境中,始终使用系统Python自带的版本,因此报错时会显示系统路径。而aiohttp的路径是虚拟环境的,说明虚拟环境已经正常生效,无需担心。
调整任务await时机
将await asyncio.gather(*async_tasks)移动到while循环内部、session上下文管理器的范围内,确保每一页的任务都在session关闭前完成:
async def get_data_from_habr() -> None: """Main function, adding info from all pages to database.""" p = 1 async with aiohttp.ClientSession() as session: while True: url = f"https://freelance.habr.com/tasks?categories=development_bots&page={p}" headers = {"user-agent": UserAgent().random} async with session.get(url=url, headers=headers) as res: src = await res.text() soup = BeautifulSoup(src, "lxml") if soup.find(class_="empty-block__title"): break orders = soup.find_all(class_="task__title") async_tasks = [] for order in orders: order_url = "https://freelance.habr.com" + order.find("a")["href"] async_tasks.append( asyncio.create_task(get_data_from_habr_order_page(order_url, session)) ) # 在session上下文内await当前页的所有任务 await asyncio.gather(*async_tasks) p += 1
优化User-Agent设置
避免每次请求都生成随机User-Agent,在创建session时统一设置一次,减少被反爬的概率:
async def get_data_from_habr() -> None: """Main function, adding info from all pages to database.""" p = 1 # 初始化session时设置默认headers async with aiohttp.ClientSession(headers={"user-agent": UserAgent().random}) as session: while True: url = f"https://freelance.habr.com/tasks?categories=development_bots&page={p}" # 无需再传headers,session会自动使用默认值 async with session.get(url=url) as res: src = await res.text() # 后续代码保持不变 soup = BeautifulSoup(src, "lxml") if soup.find(class_="empty-block__title"): break # ...
内容的提问来源于stack exchange,提问作者JFMAN12

