You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在禁用Selenium的前提下从Hiperlibertad提取全部分类/子分类?

解决方案:通过API获取全部分类 + 复用现有商品爬取逻辑

核心思路

该电商网站的分类与商品数据均通过API提供,无需依赖Playwright渲染JS。先调用分类树API获取所有一级、二级分类的路径,再遍历每个分类路径,复用你已有的商品分页爬取逻辑即可。

修改后的完整代码

import scrapy
import json

class HiperSpider(scrapy.Spider):
    name = "hiper"
    sc_value = '1'  # 门店标识,可按需修改
    # 分类树API,根分类ID为2(对应网站首页主菜单)
    category_api_url = 'https://www.hiperlibertad.com.ar/api/catalog_system/pub/category/tree/2'

    def __init__(self, sc_value='1', *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.sc_value = sc_value

    def start_requests(self):
        # 先请求分类树API,获取所有分类路径
        yield scrapy.Request(self.category_api_url, callback=self.parse_categories)

    def parse_categories(self, response):
        category_tree = json.loads(response.text)
        # 遍历一级分类
        for main_category in category_tree:
            # 遍历二级分类(子分类)
            for sub_category in main_category.get('children', []):
                # 构造分类的商品搜索API路径,格式如"hogar/muebles-de-interior"
                category_path = f"{main_category['url'].strip('/')}/{sub_category['url'].strip('/')}"
                # 构造第一页商品请求URL
                first_page_url = (
                    f"https://www.hiperlibertad.com.ar/api/catalog_system/pub/products/search/{category_path}"
                    f"?O=OrderByTopSaleDESC&_from=0&_to=23&sc={self.sc_value}"
                )
                # 传递分类层级名称到商品解析函数
                yield scrapy.Request(
                    first_page_url,
                    callback=self.parse_products,
                    meta={'category_name': f"{main_category['name']} > {sub_category['name']}"}
                )

    def parse_products(self, response):
        data = json.loads(response.text)
        category_name = response.meta['category_name']
        
        for product in data:
            name = product['productName']
            regular_price = product['items'][0]['sellers'][0]['commertialOffer']['Price']
            promotional_price = product['items'][0]['sellers'][0]['commertialOffer']['ListPrice']
            sku = product['productId']

            yield {
                'name': name,
                'regular_price': regular_price,
                'promotional_price': promotional_price,
                'category': category_name,
                'sku': sku
            }

        # 处理分页逻辑,复用原有逻辑
        has_more_products = len(data) > 0
        if has_more_products:
            current_from = int(response.url.split('_from=')[1].split('&')[0])
            current_to = int(response.url.split('_to=')[1].split('&')[0])
            next_from = current_from + 24
            next_to = current_to + 24
            # 构造下一页URL
            next_page_url = response.url.replace(f'_from={current_from}&_to={current_to}', f'_from={next_from}&_to={next_to}')
            yield scrapy.Request(
                next_page_url,
                callback=self.parse_products,
                meta={'category_name': category_name}
            )

关键说明

  • 分类API获取:通过/api/catalog_system/pub/category/tree/2获取完整分类树,根分类ID2对应网站首页主菜单,返回的JSON包含所有一级、二级分类的URL路径与名称。
  • 分类路径拼接:将一级、二级分类的URL路径拼接成商品搜索API所需格式(如hogar/muebles-de-interior)。
  • 分页逻辑复用:保留原有分页判断与URL构造逻辑,通过meta传递分类名称,确保每条商品数据关联正确的分类层级。
  • 无需JS渲染工具:所有数据均通过直接调用API获取,避免JS渲染复杂度,爬取效率更高。

可调整点

  • 若需爬取三级及以上分类,可在parse_categories中递归遍历children字段。
  • 若sc_value需动态获取,可先请求网站首页,从页面script标签或Cookie中提取对应值。

内容的提问来源于stack exchange,提问作者Kiko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 18:23:11