Scrapy爬虫无法自动翻页问题排查(附实现代码)
问题根因
你的判断完全正确:这个站点的分页逻辑是基于POST表单请求实现的,所有分页指向同一个index.php路径,靠表单里的rec_start(结果偏移量)、rec_dur(每页条数)参数区分不同页码,普通GET请求拿不到分页内容。
你现在用的CrawlSpider+LinkExtractor方案天生不适用这类场景:
- LinkExtractor只会提取a标签的href属性构造GET请求,不会自动携带你初始搜索的表单参数
- 你代码里定义的分页参数
page_2_form_data_additions从未被实际调用,rec_start偏移量不会随页码递增 - Rule参数里的
cb_kwargs=data作用是给回调函数传关键字参数,不是给请求附加表单数据,属于参数误用 - 全局字典
classDict在Scrapy异步调度逻辑下会出现数据竞争,不适合存储爬取结果 - 请求头里的Referer字段拼写错误,把域名
pisa.ucsc.edu写成了pisa.usc.edu,可能被站点拦截 - 重写的
start_requests只处理了第一页的POST请求,后续分页请求没有沿用POST逻辑和对应参数
修正方案
放弃CrawlSpider的自动爬取规则,改用基础Spider手动构造每一页的FormRequest即可,核心逻辑是每解析完一页,就根据当前页的条目数判断是否存在下一页,累加偏移量后构造新的POST请求:
import scrapy import re from scrapy import Spider # 修正请求头:移除重复的Content-Type字段,修正Referer拼写错误 all_class_headers = { 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Content-Type': 'application/x-www-form-urlencoded', 'Origin': 'https://pisa.ucsc.edu', 'Accept-Language': 'en-us', 'Host': 'pisa.ucsc.edu', 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/14.1.1 Safari/605.1.15', 'Referer': 'https://pisa.ucsc.edu/class_search/', 'Accept-Encoding': 'gzip, deflate, br', 'Connection': 'keep-alive', } # 基础表单参数,所有分页通用 base_form_data = { 'action': 'results', 'binds[:term]': '2228', 'binds[:reg_status]': 'all', 'binds[:subject]': '', 'binds[:catalog_nbr_op]': '=', 'binds[:catalog_nbr]': '', 'binds[:title]': '', 'binds[:instr_name_op]': '=', 'binds[:instructor]': '', 'binds[:ge]': '', 'binds[:crse_units_op]': '=', 'binds[:crse_units_from]': '', 'binds[:crse_units_to]': '', 'binds[:crse_units_exact]': '', 'binds[:days]': '', 'binds[:times]': '', 'binds[:acad_career]': '', 'binds[:asynch]': 'A', 'binds[:hybrid]': 'H', 'binds[:synch]': 'S', 'binds[:person]': 'P', 'rec_dur': '25' } class ClassSpider(Spider): name = "classes" allowed_domains = ['pisa.ucsc.edu'] def start_requests(self): # 构造第一页请求,偏移量从0开始 first_page_form = base_form_data.copy() first_page_form['rec_start'] = '0' yield scrapy.FormRequest( url='https://pisa.ucsc.edu/class_search/index.php', headers=all_class_headers, formdata=first_page_form, callback=self.parse_item, meta={'current_offset': 0} ) def parse_item(self, response): current_offset = response.meta['current_offset'] course_rows = response.xpath('//div[contains(@id, "rowpanel_")]') # 解析当前页课程数据 for row in course_rows: class_name = row.xpath('.//h2//a/text()').re(r'(?i)(\w+\s\w+)+\s-\s\w+\xa0+([\w\s]+\b)') professor = row.xpath('(.//div[@class="panel-body"]//div)[3]/text()').get().strip() class_number = row.xpath('(.//div[@class="panel-body"]//div)[2]/a/text()').get().strip() class_time = row.xpath('(.//div[@class="panel-body"]//div[@class="col-xs-6 col-sm-6"])[2]/text()').get().strip() location = row.xpath('(.//div[@class="panel-body"]//div[@class="col-xs-6 col-sm-6"])[1]/text()').get().strip() attendance_type = row.xpath('(.//div[@class="panel-body"]//div[@class="col-xs-6 col-sm-3 hide-print"])[3]/b/text()').get().strip() yield { 'class_number': class_number, 'professor': professor, 'class_name': class_name, 'time': class_time, 'location': location, 'online_or_in_person': attendance_type } # 当前页条目数等于每页25条说明还有下一页,累加偏移量构造下一页请求 if len(course_rows) == 25: next_offset = current_offset + 25 next_page_form = base_form_data.copy() next_page_form['rec_start'] = str(next_offset) yield scrapy.FormRequest( url='https://pisa.ucsc.edu/class_search/index.php', headers=all_class_headers, formdata=next_page_form, callback=self.parse_item, meta={'current_offset': next_offset} )
调试时可以打开浏览器开发者工具,手动点击分页按钮,核对网络请求里的Form Data参数和你构造的是否一致,不需要依赖页面上的分页按钮链接。
内容的提问来源于stack exchange,提问作者John Jacob
相关产品推荐
相关产品推荐

