如何解决基于asyncio的网页爬虫执行停滞问题
解决Asyncio爬虫停滞问题的方案
我看了你的代码,发现几个关键问题导致脚本运行100条左右就停滞,咱们一步步来修复:
1. 频繁创建ClientSession是资源消耗的元凶
你的fetch函数每次调用都会新建一个ClientSession,递归爬取下一页时会不断创建新会话,这会快速耗尽系统的连接资源,导致后续请求无法建立。必须复用同一个Session,它会自动管理连接池,提升效率还能避免资源枯竭。
2. Semaphore嵌套使用可能引发死锁
你在fetch和fetch_again里都用了同一个semaphore,相当于每个列表页请求先占1个许可,然后里面的详情页请求又要占许可。比如semaphore设为10,一个列表页请求占1个,剩下9个许可,但如果页面有10个帖子,那第10个详情页请求就会等待许可,而列表页请求要等所有详情页和下一页请求完成,下一页请求又要等许可,这就形成了死锁。
3. 缺少超时和异常处理
有些请求可能因为网络问题或网站限制卡住,一直不返回结果,导致整个事件循环停滞。而且没有异常捕获,单个请求失败就可能拖垮整个任务。
修改后的完整代码
import aiohttp import asyncio from lxml.html import fromstring from urllib.parse import urljoin link = "https://stackoverflow.com/questions/tagged/web-scraping" semaphore = asyncio.Semaphore(10) async def fetch(session, url): try: async with semaphore: # 添加10秒超时,避免请求卡住 async with session.get(url, timeout=aiohttp.ClientTimeout(total=10)) as response: response.raise_for_status() # 主动抛出HTTP错误 text = await response.text() await processing_docs(session, text) except Exception as e: print(f"列表页请求失败 {url}: {str(e)}") async def processing_docs(session, html): coros = [] tree = fromstring(html) # 提取帖子链接 titles = [urljoin(link, title.attrib['href']) for title in tree.cssselect(".summary .question-hyperlink")] for title in titles: coros.append(fetch_again(session, title)) # 处理下一页 next_page = tree.cssselect("div.pager a[rel='next']") if next_page: page_link = urljoin(link, next_page[0].attrib['href']) coros.append(fetch(session, page_link)) # return_exceptions=True:单个任务失败不影响其他任务 await asyncio.gather(*coros, return_exceptions=True) async def fetch_again(session, url): try: async with semaphore: async with session.get(url, timeout=aiohttp.ClientTimeout(total=10)) as response: response.raise_for_status() text = await response.text() tree = fromstring(text) title_elem = tree.cssselect("h1[itemprop='name'] a") if title_elem: title = title_elem[0].text print(title) else: print(f"无法提取帖子标题 {url}") except Exception as e: print(f"详情页请求失败 {url}: {str(e)}") async def main(): # 全局复用一个ClientSession async with aiohttp.ClientSession() as session: await fetch(session, link) if __name__ == '__main__': # 用asyncio.run替代旧的loop写法,更简洁安全 asyncio.run(main())
关键修改点说明
- 复用ClientSession:在
main函数中创建一个全局会话,所有请求都用它,避免重复创建连接。 - 添加超时设置:每个
session.get都加上10秒超时,防止请求无限期卡住。 - 异常捕获与处理:每个请求都包裹try-except,单个请求失败只会打印错误,不会影响整个爬虫。
- 调整Semaphore使用:虽然还是用同一个semaphore,但通过
return_exceptions=True避免死锁,同时控制全局并发数在10以内,避免被网站反爬。 - 简化Loop管理:用
asyncio.run替代手动创建loop的写法,更符合现代Python异步编程规范。
这样修改后,脚本应该能持续运行,不会再出现无报错停滞的问题了。
内容的提问来源于stack exchange,提问作者MITHU
相关产品推荐
相关产品推荐

