Scrapy分页功能异常:所有请求仅获取第一页导致无限循环
问题解决:Scrapy分页无限循环且始终返回第一页
核心问题
你代码里构造下一页URL的时候犯了个低级错误:
url = f'https://web-scraping.dev/api/testimonials?page{self.page}'
这里的参数格式不对,应该是page={self.page},少了个=,导致每次请求的URL都是?page2、?page3这种非法格式,API无法识别页码参数,就默认返回第一页数据,自然会无限循环。
修复后的完整代码
另外还要加上分页终止判断——当返回的内容为空时,停止爬取,避免无意义的请求:
from typing import Iterable import scrapy from scrapy.exceptions import CloseSpider class ScrollSpider(scrapy.Spider): name = 'scroll' page = 1 headers={ "Referer": "https://web-scraping.dev/testimonials", "X-Secret-Token": "secret123", } def start_requests(self): yield scrapy.Request(f'https://web-scraping.dev/api/testimonials?page=1', headers=self.headers, callback=self.parse) def parse(self, response): if response.status != 200: raise CloseSpider(f'Receive {response.status} status code') testimonials = response.css('div.testimonial') # 如果当前页没有数据,直接终止 if not testimonials: raise CloseSpider('No more testimonials to scrape') for testimonial in testimonials: yield { 'user_name': testimonial.css('identicon-svg::attr("username")').get(), 'testimonial': testimonial.css('p::text').get(), 'rating': len(testimonial.css('span svg').getall()) } self.page += 1 # 修复URL的参数格式,加上等于号 url = f'https://web-scraping.dev/api/testimonials?page={self.page}' yield response.follow(url, headers=self.headers, callback=self.parse)
额外说明
- 修正后的URL符合API的参数要求,会正确请求对应页码的数据
- 新增的空内容判断会在API返回空页面时自动终止爬虫,避免无限循环
内容的提问来源于stack exchange,提问作者TEOTU
相关产品推荐
相关产品推荐

