You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Playwright做JS渲染时Scrapy回调函数未执行问题

Scrapy+Playwright爬虫else分支回调未触发的解决方法

问题根源

  1. 爬虫类未继承核心父类
    你定义的Spider类没有继承scrapy.Spider,这会导致Scrapy无法正确识别爬虫实例,进而无法调度后续的回调函数,这是parse函数不执行的核心原因。

  2. 重复请求被去重机制拦截
    在else分支中,你发起了当前页面URL的请求,但这个URL已经被之前的parse_urls处理过,Scrapy默认的去重机制会直接忽略该请求,导致parse函数不会被触发。

修复方案

1. 修正爬虫类继承关系

确保你的爬虫类继承scrapy.Spider,这是Scrapy爬虫的基础要求。

2. 避免重复请求(两种可选方式)

  • 方式一:直接调用回调函数(推荐)
    在else分支中直接调用parse函数处理当前响应,无需重新发起请求,既高效又避免去重问题。
  • 方式二:跳过去重检查
    如果必须重新发起请求,在scrapy.Request中添加dont_filter=True参数,让Scrapy跳过对该URL的去重验证。

修复后的完整代码

import scrapy
from scrapy_playwright.page import PageMethod

class QuotesSpider(scrapy.Spider):
    name = "quotes"
    allowed_domains = ['quotes.toscrape.com']
  
    def start_requests(self):
        yield scrapy.Request(
            url='https://quotes.toscrape.com/js/', 
            callback=self.parse_urls, 
            meta=dict(
                playwright=True, 
                playwright_include_page=True,
                playwright_page_methods=[
                    PageMethod('wait_for_selector', 'body > div > nav > ul > li > a')
                ],
            )
        )
    

    async def parse_urls(self, response):
        page = response.meta['playwright_page']
        await page.close()
        
        next_page_url = response.xpath('//li[@class="next"]/a/@href').get()

        if next_page_url:
            print("Inside if block")
            url = 'https://quotes.toscrape.com' + next_page_url
            yield scrapy.Request(
                url=url,
                callback=self.parse_urls,
                meta=dict(
                    playwright=True,
                    playwright_include_page=True,
                    playwright_page_methods=[
                        PageMethod('wait_for_selector', 'body > div > div.quote')
                    ]
                )
            )
        else:
            print("Next page link not found")
            # 直接调用parse处理当前响应,无需重新发请求
            await self.parse(response)
            # 若需重新请求,取消下方注释并启用
            # yield scrapy.Request(
            #     url=response.request.url, 
            #     callback=self.parse, 
            #     meta=dict(
            #         playwright=True,
            #         playwright_include_page=True,
            #         playwright_page_methods=[
            #             PageMethod('wait_for_selector', 'body > div > div.quote')
            #         ]
            #     ),
            #     dont_filter=True
            # )


    async def parse(self, response):
        # 注意:如果是直接调用parse,无需重复关闭page(之前已关闭)
        # 若使用重新请求的方式,需取消下方注释
        # page = response.meta['playwright_page']
        # await page.close()
        print("Function has been called, because next page link not found")

验证效果

运行修复后的代码,日志会在Next page link not found之后,紧接着打印Function has been called, because next page link not found,说明parse函数已正常触发。

内容的提问来源于stack exchange,提问作者J775

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:43:15