You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于计数器的Scrapy品牌分页功能失效问题排查

问题分析与修复方案

你的代码主要存在几个核心问题,直接导致分页失效、品牌遍历逻辑混乱:

  1. 计数器重置问题:cgroup和cbrand在parse方法内初始化,每次调用parse(包括分页请求)都会重置为1,品牌遍历根本无法持续推进。
  2. 节点遍历错误:num_group = response.xpath(...).get()获取的是单个节点的字符串形式,for m in num_group会遍历字符而非节点,完全无法正确获取品牌链接。
  3. 页面逻辑混淆:请求品牌页面后,直接处理初始页面的商品数据,而非品牌页面的内容;同时分页请求指向的是初始页面的下一页,而非当前品牌的分页。
  4. 分页回调错误:品牌页面的分页请求回调到parse,但parse的逻辑是处理品牌分组,导致分页页面被错误解析。

修复后的代码

import scrapy

class MlSpider(scrapy.Spider):
    name = "ml"

    def start_requests(self):
        yield scrapy.Request('https://lista.mercadolivre.com.br/produtos-cabelo')

    def parse(self, response):
        # 遍历所有品牌分组
        for group in response.xpath('//div[@class="ui-search-search-modal-filter-group"]'):
            # 提取当前分组下的所有品牌链接
            brand_links = group.xpath('.//a[@class="ui-search-search-modal-filter ui-search-link"]/@href').getall()
            for link in brand_links:
                # 请求每个品牌页面,用专门的回调处理商品数据
                yield scrapy.Request(url=link, callback=self.parse_brand_page)

    def parse_brand_page(self, response):
        # 提取当前品牌页面的商品信息
        for item in response.xpath('.//div[@class="ui-search-result__content"]'):
            marca = item.xpath('.//span[@class="ui-search-item__brand-discoverability ui-search-item__group__element"]/text()').get()
            title = item.xpath('.//h2/text()').get()
            real = item.xpath('.//span[@class="andes-money-amount ui-search-price__part ui-search-price__part--medium andes-money-amount--cents-superscript"]//span[@class="andes-money-amount__fraction"]/text()').get()
            centavo = item.xpath('.//span[@class="andes-money-amount ui-search-price__part ui-search-price__part--medium andes-money-amount--cents-superscript"]//span[@class="andes-money-amount__cents andes-money-amount__cents--superscript-24"]/text()').get()
            # 处理分数字段为空的情况
            value = f'R$ {real},{centavo}' if centavo else f'R$ {real}'

            yield {
                'marca': marca,
                'title': title,
                'value': value,
                'link': item.xpath('.//a/@href').get()
            }

        # 处理当前品牌的分页逻辑
        next_page = response.xpath('//a[contains(@title,"Seguinte")]/@href').get()
        if next_page:
            yield scrapy.Request(url=next_page, callback=self.parse_brand_page)

关键修复点说明

  • 拆分回调函数:用parse专门处理品牌分组和品牌链接抓取,parse_brand_page专注于单个品牌的商品提取与分页,逻辑清晰不混淆。
  • 替换计数器逻辑:通过xpath(...).getall()直接获取所有品牌链接,彻底避免计数器带来的边界判断问题。
  • 修正分页回调:品牌页面的下一页请求回调到parse_brand_page,保证分页时持续提取当前品牌的商品数据。
  • 增加容错处理:对centavo字段为空的情况做了兼容,避免字符串格式化报错。

内容的提问来源于stack exchange,提问作者user23970789

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 19:52:40