You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Scrapy抓取滚动加载页面的商品数据?

问题描述

我刚学Scrapy,想抓取ZARA俄罗斯站点的女装新品数据,页面初始加载24个商品,向下滚动会加载更多,总共约334个。但我写的代码只能抓到24个,猜测需要用Selenium或Splash渲染页面、滚动到底部才能完整抓取。以下是当前代码:

import scrapy

custom_settings = { 
   'USER_AGENT': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36 OPR/92.0.0.0'
    }

class BookSpider(scrapy.Spider):
    name = 'basics2'
    api_url = 'https://www.zara.com/ru/ru/zhenshchiny-novinki-l1180.html?v1=2111785&page'
    start_urls = ['https://www.zara.com/ru/ru/zhenshchiny-novinki-l1180.html?v1=2111785&page=1']


#Def parse goes to the href of every product 

    def parse(self, response):
        for link in response.xpath("//div[@class='product-grid-product-info__main-info']//a"):
            yield response.follow(link, callback=self.parse_book)
        for link in response.xpath("//ul[@class='carousel__items']//li[@class='product-grid-product _product product-grid-product--ZOOM1-columns product-grid-product--0th-column']//a"):
            yield response.follow(link, callback=self.parse_book)    
        for link in response.xpath("//ul[@class='carousel__items']//li[@class='product-grid-product _product product-grid-product--ZOOM1-columns product-grid-product--1th-column']//a"):
            yield response.follow(link, callback=self.parse_book)   
        for link in response.xpath("//ul[@class='carousel__items']//li[@class='product-grid-product _product product-grid-product--ZOOM1-columns product-grid-product--th-column']//a"):
            yield response.follow(link, callback=self.parse_book)   
        for link in response.xpath("//ul[@class='carousel__items']//li[@class='product-grid-product _product carousel__item product-grid-product--ZOOM1-columns product-grid-product--0th-column']//a"):
            yield response.follow(link, callback=self.parse_book) 
        for link in response.xpath("//ul[@class='product-grid-product-info__main-info']//a"):
            yield response.follow(link, callback=self.parse_book) 


#def parse-book gets all the information inside each product
    def parse_book(self, response):
        yield{
            'title' : response.xpath("//div[@class='product-detail-info__header']/h1/text()").get(),
            'normal_price' : response.xpath("//div[@class='money-amount price-formatted__price-amount']//span//text()").get(),
            'discounted_price'  : response.xpath("(//span[@class='price__amount price__amount--on-sale price-current--with-background']//div[@class='money-amount price-formatted__price-amount']//span)[1]").get(),
            'Reference' : response.xpath("//div[@class='product-detail-color-selector product-detail-info__color-selector']//p[@class='product-detail-selected-color product-detail-color-selector__selected-color-name']//text()").get(),
            'Description'  : response.xpath("//div[@class='expandable-text__inner-content']//p//text()").get(),
            'Image' : response.xpath("//picture[@class='media-image']//source//@srcset").extract(),
            'item_url' : response.url,
            # 'User-Agent': response.request.headers['User-Agent']
    }

解决方案

方案一:抓取AJAX接口数据(推荐,效率更高)

这类滚动加载页面通常通过后台接口返回商品数据,直接爬接口比渲染页面更高效,还能避免反爬限制。

步骤:

  1. 打开浏览器开发者工具(F12),切换到「Network > XHR」标签
  2. 向下滚动页面,观察新出现的请求,找到返回商品列表的JSON接口(通常包含product、list等关键词)
  3. 分析接口的分页参数(比如page、limit),构造完整请求URL

示例代码:

import scrapy
import json

custom_settings = {
    'USER_AGENT': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36 OPR/92.0.0.0'
}

class ZaraSpider(scrapy.Spider):
    name = 'zara_spider'
    # 替换为你找到的实际接口URL,这里是示例格式
    base_api_url = 'https://www.zara.com/ru/ru/api/v1/categories/1180/products?page={}&limit=24'
    total_pages = 15  # 334/24≈14,取15确保覆盖所有商品

    def start_requests(self):
        for page in range(1, self.total_pages + 1):
            yield scrapy.Request(
                url=self.base_api_url.format(page),
                callback=self.parse_api,
                headers={
                    'Accept': 'application/json',
                    'Referer': 'https://www.zara.com/ru/ru/zhenshchiny-novinki-l1180.html?v1=2111785'
                }
            )

    def parse_api(self, response):
        data = json.loads(response.text)
        for product in data.get('products', []):
            # 从JSON中提取商品详情页链接
            product_url = f"https://www.zara.com/ru/ru{product['url']}"
            yield scrapy.Request(product_url, callback=self.parse_book)

    def parse_book(self, response):
        # 保留原有解析逻辑
        yield {
            'title': response.xpath("//div[@class='product-detail-info__header']/h1/text()").get(),
            'normal_price': response.xpath("//div[@class='money-amount price-formatted__price-amount']//span//text()").get(),
            'discounted_price': response.xpath("(//span[@class='price__amount price__amount--on-sale price-current--with-background']//div[@class='money-amount price-formatted__price-amount']//span)[1]").get(),
            'Reference': response.xpath("//div[@class='product-detail-color-selector product-detail-info__color-selector']//p[@class='product-detail-selected-color product-detail-color-selector__selected-color-name']//text()").get(),
            'Description': response.xpath("//div[@class='expandable-text__inner-content']//p//text()").get(),
            'Image': response.xpath("//picture[@class='media-image']//source//@srcset").extract(),
            'item_url': response.url
        }

方案二:用Selenium配合Scrapy渲染页面

如果找不到接口,就用Selenium模拟浏览器滚动,加载所有商品后再提取链接。

前置准备:

先安装依赖:

pip install selenium

并下载对应浏览器的驱动(比如ChromeDriver),确保驱动版本和浏览器版本匹配。

示例代码:

import scrapy
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

custom_settings = {
    'USER_AGENT': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/106.0.0.0 Safari/537.36 OPR/92.0.0.0'
}

class ZaraSeleniumSpider(scrapy.Spider):
    name = 'zara_selenium_spider'
    start_urls = ['https://www.zara.com/ru/ru/zhenshchiny-novinki-l1180.html?v1=2111785']

    def __init__(self):
        # 配置Chrome无头模式(可选,隐藏浏览器窗口)
        chrome_options = Options()
        chrome_options.add_argument('--headless=new')
        self.driver = webdriver.Chrome(options=chrome_options)

    def parse(self, response):
        self.driver.get(response.url)
        time.sleep(2)  # 等待初始页面加载

        # 模拟滚动到底部,直到所有商品加载完成
        last_scroll_height = self.driver.execute_script("return document.body.scrollHeight")
        while True:
            # 滚动到底部
            self.driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
            time.sleep(3)  # 等待新商品加载
            # 检查是否滚动到了底部
            new_scroll_height = self.driver.execute_script("return document.body.scrollHeight")
            if new_scroll_height == last_scroll_height:
                break
            last_scroll_height = new_scroll_height

        # 获取加载完成后的页面源码,用Scrapy解析
        page_source = self.driver.page_source
        selector = scrapy.Selector(text=page_source)

        # 简化XPath,提取所有商品链接(避免重复)
        product_links = selector.xpath("//a[contains(@href, '/ru/ru/') and contains(@class, 'product-grid-product-info__name')]/@href").extract()
        for link in product_links:
            yield response.follow(link, callback=self.parse_book)

        self.driver.quit()

    def parse_book(self, response):
        # 保留原有解析逻辑
        yield {
            'title': response.xpath("//div[@class='product-detail-info__header']/h1/text()").get(),
            'normal_price': response.xpath("//div[@class='money-amount price-formatted__price-amount']//span//text()").get(),
            'discounted_price': response.xpath("(//span[@class='price__amount price__amount--on-sale price-current--with-background']//div[@class='money-amount price-formatted__price-amount']//span)[1]").get(),
            'Reference': response.xpath("//div[@class='product-detail-color-selector product-detail-info__color-selector']//p[@class='product-detail-selected-color product-detail-color-selector__selected-color-name']//text()").get(),
            'Description': response.xpath("//div[@class='expandable-text__inner-content']//p//text()").get(),
            'Image': response.xpath("//picture[@class='media-image']//source//@srcset").extract(),
            'item_url': response.url
        }

注意事项

  • 方案一的接口URL需要自行通过浏览器开发者工具查找,不同站点接口结构不同
  • 使用Selenium时,调整等待时间避免触发反爬,也可添加代理、随机UA等策略
  • 频繁请求可能被封IP,建议添加请求延迟(在custom_settings中设置DOWNLOAD_DELAY=2)

内容的提问来源于stack exchange,提问作者That Guy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 05:42:07