You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取DiscoverUni动态页面:分页导航问题求助

解决动态加载页面的Scrapy分页爬取问题

针对你遇到的动态加载分页无XHR JSON数据的情况,给你两种可行的解决思路及代码实现:

思路一:排查页面分页请求逻辑(优先尝试)

先通过浏览器开发者工具定位分页触发的真实请求:

  1. 打开目标页面,开启F12的Network面板
  2. 点击「下一页」按钮,观察Network中的请求变化:
    • 若发现GET请求URL带有page/offset等参数,直接构造分页URL请求
    • 若为POST请求,提取页面中的表单参数(如__RequestVerificationToken、分页页码等),构造FormRequest

代码示例(假设为POST分页请求)

class UnispiderSpider(scrapy.Spider):
    name = 'unispider'
    start_urls = ['https://www.discoveruni.gov.uk/course-finder/results/']
    base_url = 'https://www.discoveruni.gov.uk/course-finder/results/'
    current_page = 1

    def parse(self, response):
        # 提取课程数据(保留原有逻辑)
        course_list = response.xpath(
            '//div[@class="course-finder-results__result-accordion-body-content comparison-course-area mb-4"]')
        for course in course_list:
            courseidentifier = course.xpath('@data-courseidentifier').get()
            uniname = course.xpath('@data-uniname').get()
            uniid = course.xpath('@data-uniid').get()
            coursename = course.xpath('@data-coursename').get()
            link = course.xpath('a/@href').get()
            yield {
                'courseidentifier': courseidentifier,
                'uniname': uniname,
                'uniid': uniid,
                'coursename': coursename,
                'link': link
            }

        # 提取分页所需的验证令牌(从页面隐藏元素获取)
        token = response.xpath('//input[@name="__RequestVerificationToken"]/@value').get()
        if not token:
            return

        # 构造下一页请求
        self.current_page += 1
        # 可通过页面元素判断是否还有下一页(比如检查下一页按钮是否禁用)
        has_next = response.xpath('//button[contains(@class, "pagination__next") and not(@disabled)]')
        if has_next:
            yield scrapy.FormRequest(
                url=self.base_url,
                formdata={
                    '__RequestVerificationToken': token,
                    'page': str(self.current_page),
                    # 补充页面中其他必要的表单参数(如搜索条件等,从Network中复制)
                },
                callback=self.parse
            )

思路二:用Playwright模拟浏览器翻页(纯前端渲染场景)

如果页面完全通过JS渲染、无后台接口返回数据,直接用Scrapy结合Playwright模拟浏览器点击翻页:

步骤1:安装依赖

pip install scrapy-playwright

步骤2:修改Scrapy配置(settings.py)

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,  # 测试时可改为False查看浏览器操作
    "timeout": 30000,
}

USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"

步骤3:修改Spider代码

from scrapy_playwright.page import PageCoroutine

class UnispiderSpider(scrapy.Spider):
    name = 'unispider'
    start_urls = ['https://www.discoveruni.gov.uk/course-finder/results/']

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta={
                    "playwright": True,
                    "playwright_page_coroutines": [
                        PageCoroutine("wait_for_selector", "div.course-finder-results__result-accordion-body-content"),
                    ],
                },
                callback=self.parse
            )

    def parse(self, response):
        # 提取课程数据(保留原有逻辑)
        course_list = response.xpath(
            '//div[@class="course-finder-results__result-accordion-body-content comparison-course-area mb-4"]')
        for course in course_list:
            courseidentifier = course.xpath('@data-courseidentifier').get()
            uniname = course.xpath('@data-uniname').get()
            uniid = course.xpath('@data-uniid').get()
            coursename = course.xpath('@data-coursename').get()
            link = course.xpath('a/@href').get()
            yield {
                'courseidentifier': courseidentifier,
                'uniname': uniname,
                'uniid': uniid,
                'coursename': coursename,
                'link': link
            }

        # 模拟点击下一页按钮(需根据页面实际元素调整选择器)
        next_button_selector = 'button[contains(@class, "pagination__next")]'
        has_next = response.xpath(next_button_selector + '[not(@disabled)]')
        if has_next:
            yield scrapy.Request(
                response.url,
                meta={
                    "playwright": True,
                    "playwright_page_coroutines": [
                        PageCoroutine("click", next_button_selector),
                        PageCoroutine("wait_for_selector", "div.course-finder-results__result-accordion-body-content"),
                    ],
                    "playwright_continue_page": True,  # 复用当前浏览器页面
                },
                callback=self.parse
            )

注意事项

  • 若使用思路一,务必仔细核对Network中的请求参数,遗漏参数会导致请求失败
  • 思路二中的按钮选择器需根据页面实际HTML结构调整,可通过浏览器元素审查工具获取准确选择器
  • 避免高频请求,可在settings.py中添加DOWNLOAD_DELAY = 2降低反爬风险

内容的提问来源于stack exchange,提问作者Jason_D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 05:40:23