You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用aiohttp异步爬虫遇AttributeError: NoneType无connect属性问题求助

问题描述

我正在编写针对freelance.habr.com的异步解析器,遍历特定查询的所有任务,使用aiohttp发起请求。执行过程中仅部分任务出现错误,且每次后续启动时错误数量都会增加。另外,错误信息指向系统Python路径而非虚拟环境路径,对此感到疑惑。

代码实现

import asyncio
import aiohttp
from bs4 import BeautifulSoup
from fake_useragent import UserAgent

async def get_data_from_habr() -> None:
    """Main function, adding info from all pages to database."""
    p = 1
    async with aiohttp.ClientSession() as session:
        while True:
            url = f"https://freelance.habr.com/tasks?categories=development_bots&page={p}"
            headers = {"user-agent": UserAgent().random}

            async with session.get(url=url, headers=headers) as res:
                src = await res.text()

            soup = BeautifulSoup(src, "lxml")

            if soup.find(class_="empty-block__title"):
                break

            orders = soup.find_all(class_="task__title")
            async_tasks = []
            for order in orders:
                order_url = "https://freelance.habr.com" + order.find("a")["href"]
                async_tasks.append(
                    asyncio.create_task(get_data_from_habr_order_page(order_url, session))
                )

            p += 1

    await asyncio.gather(*async_tasks)


async def get_data_from_habr_order_page(order_url: str, session: aiohttp.ClientSession) -> None:
    """Functions for getting info from one page"""
    headers = {"user-agent": UserAgent().random}

    async with session.get(url=order_url, headers=headers, timeout=1000) as res:
        src = await res.text()

错误信息

Traceback (most recent call last):
  File "\OrderScraper\habr_scraper.py", line 92, in <module>
    asyncio.run(get_data_from_habr())
  File "\AppData\Local\Programs\Python\Python312\Lib\asyncio\runners.py", line 194, in run
    return runner.run(main)
           ^^^^^^^^^^^^^^^^
  File "\AppData\Local\Programs\Python\Python312\Lib\asyncio\runners.py", line 118, in run
    return self._loop.run_until_complete(task)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "\AppData\Local\Programs\Python\Python312\Lib\asyncio\base_events.py", line 687, in run_until_complete
    return future.result()
           ^^^^^^^^^^^^^^^
  File "\OrderScraper\habr_scraper.py", line 39, in get_data_from_habr
    await asyncio.gather(*async_tasks)
  File "\OrderScraper\habr_scraper.py", line 46, in get_data_from_habr_order_page
    async with session.get(url=order_url, headers=headers, timeout=1000) as res:
  File "\OrderScraper\order_scraper\Lib\site-packages\aiohttp\client.py", line 1353, in __aenter__
    self._resp = await self._coro
                 ^^^^^^^^^^^^^^^^
  File "\OrderScraper\order_scraper\Lib\site-packages\aiohttp\client.py", line 657, in _request
    conn = await self._connector.connect(
                 ^^^^^^^^^^^^^^^^^^^^^^^
AttributeError: 'NoneType' object has no attribute 'connect'
错误原因分析

1. 核心错误:ClientSession提前关闭

代码中await asyncio.gather(*async_tasks)被放在了async with aiohttp.ClientSession() as session:的上下文管理器外部——当while循环因遇到空页面break后,session会被自动关闭(上下文管理器退出时销毁资源),此时再去await任务,任务中调用session.get()必然失败:session内部的connector已经被置为None,所以触发AttributeError: 'NoneType' object has no attribute 'connect'。

同时,async_tasks在while循环内每次都会被重置为空列表,最后gather的只有最后一页的任务;而前面页面创建的任务没有被await,会在后台运行,当session关闭后这些任务也会因session失效报错,这就是错误数量随每次启动增加的原因(未被正确await的后台任务积累)。

2. 系统Python路径显示的原因

错误栈中出现系统Python路径是正常现象:asyncio是Python标准库,不会被安装到虚拟环境中,始终使用系统Python自带的版本,因此报错时会显示系统路径。而aiohttp的路径是虚拟环境的,说明虚拟环境已经正常生效,无需担心。

修复方案

调整任务await时机

将await asyncio.gather(*async_tasks)移动到while循环内部、session上下文管理器的范围内,确保每一页的任务都在session关闭前完成:

async def get_data_from_habr() -> None:
    """Main function, adding info from all pages to database."""
    p = 1
    async with aiohttp.ClientSession() as session:
        while True:
            url = f"https://freelance.habr.com/tasks?categories=development_bots&page={p}"
            headers = {"user-agent": UserAgent().random}

            async with session.get(url=url, headers=headers) as res:
                src = await res.text()

            soup = BeautifulSoup(src, "lxml")

            if soup.find(class_="empty-block__title"):
                break

            orders = soup.find_all(class_="task__title")
            async_tasks = []
            for order in orders:
                order_url = "https://freelance.habr.com" + order.find("a")["href"]
                async_tasks.append(
                    asyncio.create_task(get_data_from_habr_order_page(order_url, session))
                )

            # 在session上下文内await当前页的所有任务
            await asyncio.gather(*async_tasks)
            p += 1

优化User-Agent设置

避免每次请求都生成随机User-Agent,在创建session时统一设置一次,减少被反爬的概率:

async def get_data_from_habr() -> None:
    """Main function, adding info from all pages to database."""
    p = 1
    # 初始化session时设置默认headers
    async with aiohttp.ClientSession(headers={"user-agent": UserAgent().random}) as session:
        while True:
            url = f"https://freelance.habr.com/tasks?categories=development_bots&page={p}"

            # 无需再传headers,session会自动使用默认值
            async with session.get(url=url) as res:
                src = await res.text()

            # 后续代码保持不变
            soup = BeautifulSoup(src, "lxml")
            if soup.find(class_="empty-block__title"):
                break
            # ...

内容的提问来源于stack exchange,提问作者JFMAN12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 21:54:56