You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy-Playwright抓取同URL多页面网站?脚本仅获第二页数据

修复Scrapy-Playwright多页面抓取问题

你的代码仅执行了一次下一页点击操作,只处理了第二页数据,没有循环检测并抓取剩余页面。以下是修复方案:

import scrapy
from scrapy_playwright.page import PageMethod
from scrapy.crawler import CrawlerProcess


class AwesomeSpideree(scrapy.Spider):
    name = "awesome"

    def start_requests(self):
        yield scrapy.Request(
            url="https://www.cia.gov/the-world-factbook/countries/",
            callback=self.parse,
            meta=dict(
                playwright=True,
                playwright_include_page=True,
                playwright_page_methods=[
                    PageMethod("screenshot", path="step1.png", full_page=True)
                ],
            )
        )

    async def parse(self, response):
        page = response.meta["playwright_page"]
        
        # 解析当前页面的国家链接
        country_lst = response.xpath("//div[@class='col-lg-9']")
        for country in country_lst:
            yield {
                "country_link": country.xpath(".//a/@href").get()
            }
        
        # 检测下一页按钮是否可用(未被禁用)
        next_btn_selector = "//div[@class='pagination-controls col-lg-6']//span[@class='pagination__arrow-right']"
        is_next_page_available = await page.locator(next_btn_selector).is_enabled()
        
        if is_next_page_available:
            # 点击下一页并等待页面加载完成
            await page.click(next_btn_selector)
            await page.wait_for_selector("//div[@class='col-lg-9']")
            
            # 获取更新后的页面响应,继续解析下一页
            current_page_num = response.meta.get('page_num', 1)
            new_response = await page.response()
            yield scrapy.Request(
                new_response.url,
                callback=self.parse,
                meta=dict(
                    playwright=True,
                    playwright_page=page,  # 复用当前页面,避免重复创建浏览器实例
                    playwright_page_methods=[
                        PageMethod("screenshot", path=f"step_{current_page_num+1}.png", full_page=True)
                    ],
                    page_num=current_page_num + 1
                )
            )
        else:
            # 无更多页面,关闭浏览器页面
            await page.close()

关键修复说明:

  • 移除初始自动点击:把原本在playwright_page_methods中的点击操作移到parse方法,手动控制页面跳转逻辑
  • 循环检测下一页:每次解析完当前页后,检查下一页按钮是否可用,若可用则继续点击并调用parse处理新页面
  • 复用Playwright页面:通过meta传递已打开的page对象,减少浏览器资源消耗
  • 添加加载等待:点击下一页后等待页面核心元素加载,确保获取最新的页面数据

内容的提问来源于stack exchange,提问作者Kfir Ben simon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 13:05:55