You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决基于asyncio的网页爬虫执行停滞问题

解决Asyncio爬虫停滞问题的方案

我看了你的代码,发现几个关键问题导致脚本运行100条左右就停滞,咱们一步步来修复:

1. 频繁创建ClientSession是资源消耗的元凶

你的fetch函数每次调用都会新建一个ClientSession,递归爬取下一页时会不断创建新会话,这会快速耗尽系统的连接资源,导致后续请求无法建立。必须复用同一个Session,它会自动管理连接池,提升效率还能避免资源枯竭。

2. Semaphore嵌套使用可能引发死锁

你在fetch和fetch_again里都用了同一个semaphore,相当于每个列表页请求先占1个许可,然后里面的详情页请求又要占许可。比如semaphore设为10,一个列表页请求占1个,剩下9个许可,但如果页面有10个帖子,那第10个详情页请求就会等待许可,而列表页请求要等所有详情页和下一页请求完成,下一页请求又要等许可,这就形成了死锁。

3. 缺少超时和异常处理

有些请求可能因为网络问题或网站限制卡住,一直不返回结果,导致整个事件循环停滞。而且没有异常捕获,单个请求失败就可能拖垮整个任务。


修改后的完整代码

import aiohttp
import asyncio
from lxml.html import fromstring
from urllib.parse import urljoin

link = "https://stackoverflow.com/questions/tagged/web-scraping"
semaphore = asyncio.Semaphore(10)

async def fetch(session, url):
    try:
        async with semaphore:
            # 添加10秒超时,避免请求卡住
            async with session.get(url, timeout=aiohttp.ClientTimeout(total=10)) as response:
                response.raise_for_status()  # 主动抛出HTTP错误
                text = await response.text()
                await processing_docs(session, text)
    except Exception as e:
        print(f"列表页请求失败 {url}: {str(e)}")

async def processing_docs(session, html):
    coros = []
    tree = fromstring(html)
    # 提取帖子链接
    titles = [urljoin(link, title.attrib['href']) for title in tree.cssselect(".summary .question-hyperlink")]
    for title in titles:
        coros.append(fetch_again(session, title))
    # 处理下一页
    next_page = tree.cssselect("div.pager a[rel='next']")
    if next_page:
        page_link = urljoin(link, next_page[0].attrib['href'])
        coros.append(fetch(session, page_link))
    # return_exceptions=True:单个任务失败不影响其他任务
    await asyncio.gather(*coros, return_exceptions=True)

async def fetch_again(session, url):
    try:
        async with semaphore:
            async with session.get(url, timeout=aiohttp.ClientTimeout(total=10)) as response:
                response.raise_for_status()
                text = await response.text()
                tree = fromstring(text)
                title_elem = tree.cssselect("h1[itemprop='name'] a")
                if title_elem:
                    title = title_elem[0].text
                    print(title)
                else:
                    print(f"无法提取帖子标题 {url}")
    except Exception as e:
        print(f"详情页请求失败 {url}: {str(e)}")

async def main():
    # 全局复用一个ClientSession
    async with aiohttp.ClientSession() as session:
        await fetch(session, link)

if __name__ == '__main__':
    # 用asyncio.run替代旧的loop写法,更简洁安全
    asyncio.run(main())

关键修改点说明

  • 复用ClientSession:在main函数中创建一个全局会话,所有请求都用它,避免重复创建连接。
  • 添加超时设置:每个session.get都加上10秒超时,防止请求无限期卡住。
  • 异常捕获与处理:每个请求都包裹try-except,单个请求失败只会打印错误,不会影响整个爬虫。
  • 调整Semaphore使用:虽然还是用同一个semaphore,但通过return_exceptions=True避免死锁,同时控制全局并发数在10以内,避免被网站反爬。
  • 简化Loop管理:用asyncio.run替代手动创建loop的写法,更符合现代Python异步编程规范。

这样修改后,脚本应该能持续运行,不会再出现无报错停滞的问题了。

内容的提问来源于stack exchange,提问作者MITHU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 20:27:49