如何在禁用Selenium的前提下从Hiperlibertad提取全部分类/子分类?
解决方案:通过API获取全部分类 + 复用现有商品爬取逻辑
核心思路
该电商网站的分类与商品数据均通过API提供,无需依赖Playwright渲染JS。先调用分类树API获取所有一级、二级分类的路径,再遍历每个分类路径,复用你已有的商品分页爬取逻辑即可。
修改后的完整代码
import scrapy import json class HiperSpider(scrapy.Spider): name = "hiper" sc_value = '1' # 门店标识,可按需修改 # 分类树API,根分类ID为2(对应网站首页主菜单) category_api_url = 'https://www.hiperlibertad.com.ar/api/catalog_system/pub/category/tree/2' def __init__(self, sc_value='1', *args, **kwargs): super().__init__(*args, **kwargs) self.sc_value = sc_value def start_requests(self): # 先请求分类树API,获取所有分类路径 yield scrapy.Request(self.category_api_url, callback=self.parse_categories) def parse_categories(self, response): category_tree = json.loads(response.text) # 遍历一级分类 for main_category in category_tree: # 遍历二级分类(子分类) for sub_category in main_category.get('children', []): # 构造分类的商品搜索API路径,格式如"hogar/muebles-de-interior" category_path = f"{main_category['url'].strip('/')}/{sub_category['url'].strip('/')}" # 构造第一页商品请求URL first_page_url = ( f"https://www.hiperlibertad.com.ar/api/catalog_system/pub/products/search/{category_path}" f"?O=OrderByTopSaleDESC&_from=0&_to=23&sc={self.sc_value}" ) # 传递分类层级名称到商品解析函数 yield scrapy.Request( first_page_url, callback=self.parse_products, meta={'category_name': f"{main_category['name']} > {sub_category['name']}"} ) def parse_products(self, response): data = json.loads(response.text) category_name = response.meta['category_name'] for product in data: name = product['productName'] regular_price = product['items'][0]['sellers'][0]['commertialOffer']['Price'] promotional_price = product['items'][0]['sellers'][0]['commertialOffer']['ListPrice'] sku = product['productId'] yield { 'name': name, 'regular_price': regular_price, 'promotional_price': promotional_price, 'category': category_name, 'sku': sku } # 处理分页逻辑,复用原有逻辑 has_more_products = len(data) > 0 if has_more_products: current_from = int(response.url.split('_from=')[1].split('&')[0]) current_to = int(response.url.split('_to=')[1].split('&')[0]) next_from = current_from + 24 next_to = current_to + 24 # 构造下一页URL next_page_url = response.url.replace(f'_from={current_from}&_to={current_to}', f'_from={next_from}&_to={next_to}') yield scrapy.Request( next_page_url, callback=self.parse_products, meta={'category_name': category_name} )
关键说明
- 分类API获取:通过
/api/catalog_system/pub/category/tree/2获取完整分类树,根分类ID2对应网站首页主菜单,返回的JSON包含所有一级、二级分类的URL路径与名称。 - 分类路径拼接:将一级、二级分类的URL路径拼接成商品搜索API所需格式(如
hogar/muebles-de-interior)。 - 分页逻辑复用:保留原有分页判断与URL构造逻辑,通过
meta传递分类名称,确保每条商品数据关联正确的分类层级。 - 无需JS渲染工具:所有数据均通过直接调用API获取,避免JS渲染复杂度,爬取效率更高。
可调整点
- 若需爬取三级及以上分类,可在
parse_categories中递归遍历children字段。 - 若
sc_value需动态获取,可先请求网站首页,从页面script标签或Cookie中提取对应值。
内容的提问来源于stack exchange,提问作者Kiko
相关产品推荐
相关产品推荐

