Scrapy爬取无限滚动页面时分页无数据问题求助
问题分析
这个网站用的是无限滚动加载,不是传统的HTML翻页模式。你直接请求/de-de/herren?p=4这类页面时,服务器返回的HTML里根本没有商品数据——因为这些页面的商品是滚动时通过AJAX接口动态加载的,并非预渲染在HTML中。所以你的代码到第3页后,第4页的products为空,触发了if len(products) > 0:的判断,直接停止了后续请求。
解决方案
直接抓取网站的商品数据API接口,而非解析HTML页面。具体步骤:
- 打开浏览器开发者工具(F12)切换到Network标签,滚动页面观察XHR请求,会发现滚动时会调用类似
https://www.salewa.com/de-de/herren?page=4&sz=36的接口,返回JSON格式的商品数据。其中sz是每页商品数量,page是页码。 - 修改代码,直接请求这些API接口,解析返回的JSON数据。
- 循环递增页码,直到返回的商品为空时停止请求。
修改后的代码
import scrapy import json class Salewa_Spider(scrapy.Spider): name = "salewa" allowed_domains = ["salewa.com"] # 初始页码和每页商品数量 current_page = 1 per_page = 36 base_api_url = "https://www.salewa.com/de-de/herren?page={}&sz={}" def start_requests(self): # 发起第一页的API请求 yield scrapy.Request( url=self.base_api_url.format(self.current_page, self.per_page), callback=self.parse_api, headers={ # 模拟浏览器请求头,避免被拦截 "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "application/json, text/javascript, */*; q=0.01" } ) def parse_api(self, response): try: data = json.loads(response.text) except json.JSONDecodeError: self.logger.error("JSON解析失败") return # 从JSON中提取商品列表 products = data.get("products", []) for product in products: yield { "name": product.get("name", "").strip(), "price": product.get("price", {}).get("sales", "").strip(), "url": f"https://www.salewa.com{product.get('url', '')}" } # 如果当前页商品数量等于每页设定值,继续请求下一页 if len(products) == self.per_page: self.current_page += 1 yield scrapy.Request( url=self.base_api_url.format(self.current_page, self.per_page), callback=self.parse_api, headers={ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "application/json, text/javascript, */*; q=0.01" } )
关键说明
- 直接请求API接口返回的结构化JSON,比解析HTML更稳定,也能获取到所有商品数据。
- 通过判断当前页商品数量是否等于每页设定值(36),来决定是否继续请求下一页——如果最后一页商品不足36,就自动停止。
- 带上
User-Agent和Accept请求头,模拟浏览器行为,降低被网站反爬拦截的概率。
内容的提问来源于stack exchange,提问作者Vaidas
相关产品推荐
相关产品推荐

