Scrapy+Playwright爬虫运行反复卡顿问题排查与解决方案
Scrapy+Playwright爬虫运行中无规律卡顿、爬速归零问题解决方案
问题表现
Scrapy集成Playwright编写的爬虫启动后前期运行正常,运行过程中不定时出现停滞,控制台输出如下日志,显示每分钟爬取页数、采集条目数均为0:
[scrapy.extensions.logstats] INFO: Crawled 1795 pages (at 0 pages/min), scraped 1716 items (at 0 items/min)
手动终止进程重启后,爬取部分数据会再次复现相同卡顿问题。
问题对应核心爬虫代码如下:
import scrapy from scrapy.loader import ItemLoader from healthgrades.items import HealthgradesItem from scrapy_playwright.page import PageMethod # 构造请求头字典 def get_headers(s, sep=': ', strip_cookie=True, strip_cl=True, strip_headers: list = []) -> dict(): d = dict() for kv in s.split('\n'): kv = kv.strip() if kv and sep in kv: v='' k = kv.split(sep)[0] if len(kv.split(sep)) == 1: v = '' else: v = kv.split(sep)[1] if v == '\'\'': v ='' if strip_cookie and k.lower() == 'cookie': continue if strip_cl and k.lower() == 'content-length': continue if k in strip_headers: continue d[k] = v return d # 爬虫类 class DoctorSpider(scrapy.Spider): name = 'doctor' allowed_domains = ['healthgrades.com'] url = 'https://www.healthgrades.com/usearch?what=Massage%20Therapy&entityCode=PS444&where=New%20York&pageNum={}&sort.provider=bestmatch&=' # 构造模拟浏览器请求头 def start_requests(self): h = get_headers( ''' accept: */* accept-encoding: gzip, deflate, be accept-language: en-US,en;q=0.9 dnt: 1 origin: https://www.healthgrades.com referer: https://www.healthgrades.com/ sec-ch-ua: ".Not/A)Brand";v="99", "Google Chrome";v="103", "Chromium";v="103" sec-ch-ua-mobile: ?0 sec-ch-ua-platform: "Windows" sec-fetch-dest: empty sec-fetch-mode: cors sec-fetch-site: cross-site user-agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36 ''' ) for i in range(1, 6): yield scrapy.Request(self.url.format(i), headers =h, meta=dict( playwright = True, playwright_include_page = True, playwright_page_methods = [PageMethod('wait_for_selector', 'h3.card-name a')] )) def parse(self, response): for link in response.css('div h3.card-name a::attr(href)'): yield response.follow(link.get(), callback = self.parse_categories) def parse_categories(self, response): l = ItemLoader(item = HealthgradesItem(), selector = response) l.add_xpath('name', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[1]/div[1]/h1') l.add_xpath('specialty', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[1]/div[1]/div[2]/p/span[1]') l.add_xpath('practice_name', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[2]/div[1]/p') l.add_xpath('address', 'string(//*[@id="summary-section"]/div[1]/div[2]/div/div/div[2]/div[1]/address)') yield l.load_item()
问题根因
这是Scrapy+Playwright组合使用时的典型故障,三个核心问题直接导致卡顿:
- 开启
playwright_include_page = True后,所有请求生成的Playwright Page对象未手动关闭,Scrapy-Playwright不会自动回收这部分资源,叠加Chromium本身的内存占用问题,运行到上千页后内存耗尽,浏览器进程无响应,新请求无法调度 wait_for_selector方法未设置超时时间,部分页面因网络波动、反爬拦截导致目标元素一直无法加载时,请求会无限挂起,占满Scrapy固定的并发槽位后,整个爬虫没有空余资源处理新请求,直接表现为爬速归零- 详情页的
response.follow请求未携带Playwright相关配置,默认用普通HTTP请求发起访问,被站点反爬策略拦截后连接挂起,进一步占用并发槽位
修复方案
按以下步骤修改即可彻底解决卡顿问题:
- 手动回收Page资源:所有使用Playwright的请求,回调处理完成后必须手动关闭Page对象,回调方法改为异步实现。修改后的解析方法示例:
async def parse(self, response): # 列表页逻辑处理完立刻关闭页面释放内存 page = response.meta["playwright_page"] await page.close() for link in response.css('div h3.card-name a::attr(href)'): yield response.follow(link.get(), callback = self.parse_categories, meta=dict( playwright = True, playwright_include_page = True, playwright_page_methods = [PageMethod('wait_for_selector', '#summary-section', timeout=10000)] )) async def parse_categories(self, response): # 详情页逻辑处理完立刻关闭页面 page = response.meta["playwright_page"] await page.close() l = ItemLoader(item = HealthgradesItem(), selector = response) l.add_xpath('name', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[1]/div[1]/h1') l.add_xpath('specialty', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[1]/div[1]/div[2]/p/span[1]') l.add_xpath('practice_name', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[2]/div[1]/p') l.add_xpath('address', 'string(//*[@id="summary-section"]/div[1]/div[2]/div/div/div[2]/div[1]/address)') yield l.load_item() - 给页面等待逻辑加超时:所有
PageMethod的等待操作统一加上10秒超时,超时后自动抛出异常,触发Scrapy重试逻辑,避免无限等待。列表页的等待逻辑修改为:playwright_page_methods = [PageMethod('wait_for_selector', 'h3.card-name a', timeout=10000)] - 调整配置降低资源占用:在
settings.py中修改/新增以下配置:# Playwright资源占用高,并发数不要超过8 CONCURRENT_REQUESTS = 6 # 开启失败重试 RETRY_TIMES = 3 RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429] # 拦截不影响DOM渲染的非必要资源,大幅降低加载耗时和内存占用 PLAYWRIGHT_ABORT_REQUEST = lambda req: req.resource_type in ["image", "font", "media", "ping", "beacon"]
验证标准
修改完成后启动爬虫,持续运行2小时以上,控制台日志中爬取速度不会归零,无长时间无响应情况即为修复完成。
内容的提问来源于stack exchange,提问作者Shahidul Islam Pranto
相关产品推荐
相关产品推荐

