You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy+Playwright爬虫运行反复卡顿问题排查与解决方案

Scrapy+Playwright爬虫运行中无规律卡顿、爬速归零问题解决方案

问题表现

Scrapy集成Playwright编写的爬虫启动后前期运行正常,运行过程中不定时出现停滞,控制台输出如下日志,显示每分钟爬取页数、采集条目数均为0:

[scrapy.extensions.logstats] INFO: Crawled 1795 pages (at 0 pages/min), scraped 1716 items (at 0 items/min)

手动终止进程重启后,爬取部分数据会再次复现相同卡顿问题。
问题对应核心爬虫代码如下:

import scrapy
from scrapy.loader import ItemLoader
from healthgrades.items import HealthgradesItem
from scrapy_playwright.page import PageMethod 

# 构造请求头字典
def get_headers(s, sep=': ', strip_cookie=True, strip_cl=True, strip_headers: list = []) -> dict():
    d = dict()
    for kv in s.split('\n'):
        kv = kv.strip()
        if kv and sep in kv:
            v=''
            k = kv.split(sep)[0]
            if len(kv.split(sep)) == 1:
                v = ''
            else:
                v = kv.split(sep)[1]
            if v == '\'\'':
                v =''
            if strip_cookie and k.lower() == 'cookie': continue
            if strip_cl and k.lower() == 'content-length': continue
            if k in strip_headers: continue
            d[k] = v
    return d

# 爬虫类
class DoctorSpider(scrapy.Spider):
    name = 'doctor'
    allowed_domains = ['healthgrades.com']
    url = 'https://www.healthgrades.com/usearch?what=Massage%20Therapy&entityCode=PS444&where=New%20York&pageNum={}&sort.provider=bestmatch&='

    # 构造模拟浏览器请求头
    def start_requests(self):
        h = get_headers(
            '''
            accept: */*
            accept-encoding: gzip, deflate, be
            accept-language: en-US,en;q=0.9
            dnt: 1
            origin: https://www.healthgrades.com
            referer: https://www.healthgrades.com/
            sec-ch-ua: ".Not/A)Brand";v="99", "Google Chrome";v="103", "Chromium";v="103"
            sec-ch-ua-mobile: ?0
            sec-ch-ua-platform: "Windows"
            sec-fetch-dest: empty
            sec-fetch-mode: cors
            sec-fetch-site: cross-site
            user-agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36
            '''
        )

        for i in range(1, 6):
            yield scrapy.Request(self.url.format(i), headers =h, meta=dict(
                playwright = True,
                playwright_include_page = True,
                playwright_page_methods = [PageMethod('wait_for_selector', 'h3.card-name a')]
            )) 

    def parse(self, response):
        for link in response.css('div h3.card-name a::attr(href)'):
            yield response.follow(link.get(), callback = self.parse_categories)
        
    def parse_categories(self, response):
        l = ItemLoader(item  = HealthgradesItem(), selector = response)

        l.add_xpath('name', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[1]/div[1]/h1')
        l.add_xpath('specialty', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[1]/div[1]/div[2]/p/span[1]')
        l.add_xpath('practice_name', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[2]/div[1]/p')
        l.add_xpath('address', 'string(//*[@id="summary-section"]/div[1]/div[2]/div/div/div[2]/div[1]/address)')

        yield l.load_item()

问题根因

这是Scrapy+Playwright组合使用时的典型故障,三个核心问题直接导致卡顿:

  • 开启playwright_include_page = True后,所有请求生成的Playwright Page对象未手动关闭,Scrapy-Playwright不会自动回收这部分资源,叠加Chromium本身的内存占用问题,运行到上千页后内存耗尽,浏览器进程无响应,新请求无法调度
  • wait_for_selector方法未设置超时时间,部分页面因网络波动、反爬拦截导致目标元素一直无法加载时,请求会无限挂起,占满Scrapy固定的并发槽位后,整个爬虫没有空余资源处理新请求,直接表现为爬速归零
  • 详情页的response.follow请求未携带Playwright相关配置,默认用普通HTTP请求发起访问,被站点反爬策略拦截后连接挂起,进一步占用并发槽位

修复方案

按以下步骤修改即可彻底解决卡顿问题:

  • 手动回收Page资源:所有使用Playwright的请求,回调处理完成后必须手动关闭Page对象,回调方法改为异步实现。修改后的解析方法示例:
    async def parse(self, response):
        # 列表页逻辑处理完立刻关闭页面释放内存
        page = response.meta["playwright_page"]
        await page.close()
        for link in response.css('div h3.card-name a::attr(href)'):
            yield response.follow(link.get(), callback = self.parse_categories, meta=dict(
                playwright = True,
                playwright_include_page = True,
                playwright_page_methods = [PageMethod('wait_for_selector', '#summary-section', timeout=10000)]
            ))
    
    async def parse_categories(self, response):
        # 详情页逻辑处理完立刻关闭页面
        page = response.meta["playwright_page"]
        await page.close()
        l = ItemLoader(item  = HealthgradesItem(), selector = response)
        l.add_xpath('name', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[1]/div[1]/h1')
        l.add_xpath('specialty', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[1]/div[1]/div[2]/p/span[1]')
        l.add_xpath('practice_name', '//*[@id="summary-section"]/div[1]/div[2]/div/div/div[2]/div[1]/p')
        l.add_xpath('address', 'string(//*[@id="summary-section"]/div[1]/div[2]/div/div/div[2]/div[1]/address)')
        yield l.load_item()
    
  • 给页面等待逻辑加超时:所有PageMethod的等待操作统一加上10秒超时,超时后自动抛出异常,触发Scrapy重试逻辑,避免无限等待。列表页的等待逻辑修改为:
    playwright_page_methods = [PageMethod('wait_for_selector', 'h3.card-name a', timeout=10000)]
    
  • 调整配置降低资源占用:在settings.py中修改/新增以下配置:
    # Playwright资源占用高,并发数不要超过8
    CONCURRENT_REQUESTS = 6
    # 开启失败重试
    RETRY_TIMES = 3
    RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429]
    # 拦截不影响DOM渲染的非必要资源,大幅降低加载耗时和内存占用
    PLAYWRIGHT_ABORT_REQUEST = lambda req: req.resource_type in ["image", "font", "media", "ping", "beacon"]
    

验证标准

修改完成后启动爬虫,持续运行2小时以上,控制台日志中爬取速度不会归零,无长时间无响应情况即为修复完成。


内容的提问来源于stack exchange,提问作者Shahidul Islam Pranto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 06:06:55