You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy分页功能异常:所有请求仅获取第一页导致无限循环

问题解决:Scrapy分页无限循环且始终返回第一页

核心问题

你代码里构造下一页URL的时候犯了个低级错误:

url = f'https://web-scraping.dev/api/testimonials?page{self.page}'

这里的参数格式不对,应该是page={self.page},少了个=,导致每次请求的URL都是?page2、?page3这种非法格式,API无法识别页码参数,就默认返回第一页数据,自然会无限循环。

修复后的完整代码

另外还要加上分页终止判断——当返回的内容为空时,停止爬取,避免无意义的请求:

from typing import Iterable
import scrapy
from scrapy.exceptions import CloseSpider   

class ScrollSpider(scrapy.Spider):
    name = 'scroll'

    page = 1
    headers={
        "Referer": "https://web-scraping.dev/testimonials",
        "X-Secret-Token": "secret123",
    }

    def start_requests(self):
        yield scrapy.Request(f'https://web-scraping.dev/api/testimonials?page=1', headers=self.headers, callback=self.parse)

    
    def parse(self, response):
        if response.status != 200:
            raise CloseSpider(f'Receive {response.status} status code')
        
        testimonials = response.css('div.testimonial')
        # 如果当前页没有数据,直接终止
        if not testimonials:
            raise CloseSpider('No more testimonials to scrape')
            
        for testimonial in testimonials:
            yield {
                'user_name': testimonial.css('identicon-svg::attr("username")').get(),
                'testimonial': testimonial.css('p::text').get(),
                'rating': len(testimonial.css('span svg').getall())
            }

        self.page += 1 
        # 修复URL的参数格式,加上等于号
        url = f'https://web-scraping.dev/api/testimonials?page={self.page}'
        yield response.follow(url, headers=self.headers, callback=self.parse)

额外说明

  1. 修正后的URL符合API的参数要求,会正确请求对应页码的数据
  2. 新增的空内容判断会在API返回空页面时自动终止爬虫,避免无限循环

内容的提问来源于stack exchange,提问作者TEOTU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 01:51:06