You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫无法自动翻页问题排查(附实现代码)

问题根因

你的判断完全正确:这个站点的分页逻辑是基于POST表单请求实现的,所有分页指向同一个index.php路径,靠表单里的rec_start(结果偏移量)、rec_dur(每页条数)参数区分不同页码,普通GET请求拿不到分页内容。
你现在用的CrawlSpider+LinkExtractor方案天生不适用这类场景:

  • LinkExtractor只会提取a标签的href属性构造GET请求,不会自动携带你初始搜索的表单参数
  • 你代码里定义的分页参数page_2_form_data_additions从未被实际调用,rec_start偏移量不会随页码递增
  • Rule参数里的cb_kwargs=data作用是给回调函数传关键字参数,不是给请求附加表单数据,属于参数误用
  • 全局字典classDict在Scrapy异步调度逻辑下会出现数据竞争,不适合存储爬取结果
  • 请求头里的Referer字段拼写错误,把域名pisa.ucsc.edu写成了pisa.usc.edu,可能被站点拦截
  • 重写的start_requests只处理了第一页的POST请求,后续分页请求没有沿用POST逻辑和对应参数
修正方案

放弃CrawlSpider的自动爬取规则,改用基础Spider手动构造每一页的FormRequest即可,核心逻辑是每解析完一页,就根据当前页的条目数判断是否存在下一页,累加偏移量后构造新的POST请求:

import scrapy
import re
from scrapy import Spider

# 修正请求头:移除重复的Content-Type字段,修正Referer拼写错误
all_class_headers = {
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
    'Content-Type': 'application/x-www-form-urlencoded',
    'Origin': 'https://pisa.ucsc.edu',
    'Accept-Language': 'en-us',
    'Host': 'pisa.ucsc.edu',
    'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/14.1.1 Safari/605.1.15',
    'Referer': 'https://pisa.ucsc.edu/class_search/',
    'Accept-Encoding': 'gzip, deflate, br',
    'Connection': 'keep-alive',
}

# 基础表单参数,所有分页通用
base_form_data = {
    'action': 'results',
    'binds[:term]': '2228',
    'binds[:reg_status]': 'all',
    'binds[:subject]': '',
    'binds[:catalog_nbr_op]': '=',
    'binds[:catalog_nbr]': '',
    'binds[:title]': '',
    'binds[:instr_name_op]': '=',
    'binds[:instructor]': '',
    'binds[:ge]': '',
    'binds[:crse_units_op]': '=',
    'binds[:crse_units_from]': '',
    'binds[:crse_units_to]': '',
    'binds[:crse_units_exact]': '',
    'binds[:days]': '',
    'binds[:times]': '',
    'binds[:acad_career]': '',
    'binds[:asynch]': 'A',
    'binds[:hybrid]': 'H',
    'binds[:synch]': 'S',
    'binds[:person]': 'P',
    'rec_dur': '25'
}

class ClassSpider(Spider):
    name = "classes"
    allowed_domains = ['pisa.ucsc.edu']

    def start_requests(self):
        # 构造第一页请求,偏移量从0开始
        first_page_form = base_form_data.copy()
        first_page_form['rec_start'] = '0'
        yield scrapy.FormRequest(
            url='https://pisa.ucsc.edu/class_search/index.php',
            headers=all_class_headers,
            formdata=first_page_form,
            callback=self.parse_item,
            meta={'current_offset': 0}
        )

    def parse_item(self, response):
        current_offset = response.meta['current_offset']
        course_rows = response.xpath('//div[contains(@id, "rowpanel_")]')
        
        # 解析当前页课程数据
        for row in course_rows:
            class_name = row.xpath('.//h2//a/text()').re(r'(?i)(\w+\s\w+)+\s-\s\w+\xa0+([\w\s]+\b)')
            professor = row.xpath('(.//div[@class="panel-body"]//div)[3]/text()').get().strip()
            class_number = row.xpath('(.//div[@class="panel-body"]//div)[2]/a/text()').get().strip()
            class_time = row.xpath('(.//div[@class="panel-body"]//div[@class="col-xs-6 col-sm-6"])[2]/text()').get().strip()
            location = row.xpath('(.//div[@class="panel-body"]//div[@class="col-xs-6 col-sm-6"])[1]/text()').get().strip()
            attendance_type = row.xpath('(.//div[@class="panel-body"]//div[@class="col-xs-6 col-sm-3 hide-print"])[3]/b/text()').get().strip()
            
            yield {
                'class_number': class_number,
                'professor': professor,
                'class_name': class_name,
                'time': class_time,
                'location': location,
                'online_or_in_person': attendance_type
            }

        # 当前页条目数等于每页25条说明还有下一页,累加偏移量构造下一页请求
        if len(course_rows) == 25:
            next_offset = current_offset + 25
            next_page_form = base_form_data.copy()
            next_page_form['rec_start'] = str(next_offset)
            yield scrapy.FormRequest(
                url='https://pisa.ucsc.edu/class_search/index.php',
                headers=all_class_headers,
                formdata=next_page_form,
                callback=self.parse_item,
                meta={'current_offset': next_offset}
            )

调试时可以打开浏览器开发者工具,手动点击分页按钮,核对网络请求里的Form Data参数和你构造的是否一致,不需要依赖页面上的分页按钮链接。

内容的提问来源于stack exchange,提问作者John Jacob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 00:27:25