You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取无限滚动页面时分页无数据问题求助

问题分析

这个网站用的是无限滚动加载,不是传统的HTML翻页模式。你直接请求/de-de/herren?p=4这类页面时,服务器返回的HTML里根本没有商品数据——因为这些页面的商品是滚动时通过AJAX接口动态加载的,并非预渲染在HTML中。所以你的代码到第3页后,第4页的products为空,触发了if len(products) > 0:的判断,直接停止了后续请求。

解决方案

直接抓取网站的商品数据API接口,而非解析HTML页面。具体步骤:

  1. 打开浏览器开发者工具(F12)切换到Network标签,滚动页面观察XHR请求,会发现滚动时会调用类似https://www.salewa.com/de-de/herren?page=4&sz=36的接口,返回JSON格式的商品数据。其中sz是每页商品数量,page是页码。
  2. 修改代码,直接请求这些API接口,解析返回的JSON数据。
  3. 循环递增页码,直到返回的商品为空时停止请求。
修改后的代码
import scrapy
import json

class Salewa_Spider(scrapy.Spider):
    name = "salewa"
    allowed_domains = ["salewa.com"]
    # 初始页码和每页商品数量
    current_page = 1
    per_page = 36
    base_api_url = "https://www.salewa.com/de-de/herren?page={}&sz={}"

    def start_requests(self):
        # 发起第一页的API请求
        yield scrapy.Request(
            url=self.base_api_url.format(self.current_page, self.per_page),
            callback=self.parse_api,
            headers={
                # 模拟浏览器请求头,避免被拦截
                "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
                "Accept": "application/json, text/javascript, */*; q=0.01"
            }
        )

    def parse_api(self, response):
        try:
            data = json.loads(response.text)
        except json.JSONDecodeError:
            self.logger.error("JSON解析失败")
            return

        # 从JSON中提取商品列表
        products = data.get("products", [])
        for product in products:
            yield {
                "name": product.get("name", "").strip(),
                "price": product.get("price", {}).get("sales", "").strip(),
                "url": f"https://www.salewa.com{product.get('url', '')}"
            }

        # 如果当前页商品数量等于每页设定值,继续请求下一页
        if len(products) == self.per_page:
            self.current_page += 1
            yield scrapy.Request(
                url=self.base_api_url.format(self.current_page, self.per_page),
                callback=self.parse_api,
                headers={
                    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
                    "Accept": "application/json, text/javascript, */*; q=0.01"
                }
            )
关键说明
  • 直接请求API接口返回的结构化JSON,比解析HTML更稳定,也能获取到所有商品数据。
  • 通过判断当前页商品数量是否等于每页设定值(36),来决定是否继续请求下一页——如果最后一页商品不足36,就自动停止。
  • 带上User-Agent和Accept请求头,模拟浏览器行为,降低被网站反爬拦截的概率。

内容的提问来源于stack exchange,提问作者Vaidas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 07:35:32