基于计数器的Scrapy品牌分页功能失效问题排查
问题分析与修复方案
你的代码主要存在几个核心问题,直接导致分页失效、品牌遍历逻辑混乱:
- 计数器重置问题:
cgroup和cbrand在parse方法内初始化,每次调用parse(包括分页请求)都会重置为1,品牌遍历根本无法持续推进。 - 节点遍历错误:
num_group = response.xpath(...).get()获取的是单个节点的字符串形式,for m in num_group会遍历字符而非节点,完全无法正确获取品牌链接。 - 页面逻辑混淆:请求品牌页面后,直接处理初始页面的商品数据,而非品牌页面的内容;同时分页请求指向的是初始页面的下一页,而非当前品牌的分页。
- 分页回调错误:品牌页面的分页请求回调到
parse,但parse的逻辑是处理品牌分组,导致分页页面被错误解析。
修复后的代码
import scrapy class MlSpider(scrapy.Spider): name = "ml" def start_requests(self): yield scrapy.Request('https://lista.mercadolivre.com.br/produtos-cabelo') def parse(self, response): # 遍历所有品牌分组 for group in response.xpath('//div[@class="ui-search-search-modal-filter-group"]'): # 提取当前分组下的所有品牌链接 brand_links = group.xpath('.//a[@class="ui-search-search-modal-filter ui-search-link"]/@href').getall() for link in brand_links: # 请求每个品牌页面,用专门的回调处理商品数据 yield scrapy.Request(url=link, callback=self.parse_brand_page) def parse_brand_page(self, response): # 提取当前品牌页面的商品信息 for item in response.xpath('.//div[@class="ui-search-result__content"]'): marca = item.xpath('.//span[@class="ui-search-item__brand-discoverability ui-search-item__group__element"]/text()').get() title = item.xpath('.//h2/text()').get() real = item.xpath('.//span[@class="andes-money-amount ui-search-price__part ui-search-price__part--medium andes-money-amount--cents-superscript"]//span[@class="andes-money-amount__fraction"]/text()').get() centavo = item.xpath('.//span[@class="andes-money-amount ui-search-price__part ui-search-price__part--medium andes-money-amount--cents-superscript"]//span[@class="andes-money-amount__cents andes-money-amount__cents--superscript-24"]/text()').get() # 处理分数字段为空的情况 value = f'R$ {real},{centavo}' if centavo else f'R$ {real}' yield { 'marca': marca, 'title': title, 'value': value, 'link': item.xpath('.//a/@href').get() } # 处理当前品牌的分页逻辑 next_page = response.xpath('//a[contains(@title,"Seguinte")]/@href').get() if next_page: yield scrapy.Request(url=next_page, callback=self.parse_brand_page)
关键修复点说明
- 拆分回调函数:用
parse专门处理品牌分组和品牌链接抓取,parse_brand_page专注于单个品牌的商品提取与分页,逻辑清晰不混淆。 - 替换计数器逻辑:通过
xpath(...).getall()直接获取所有品牌链接,彻底避免计数器带来的边界判断问题。 - 修正分页回调:品牌页面的下一页请求回调到
parse_brand_page,保证分页时持续提取当前品牌的商品数据。 - 增加容错处理:对
centavo字段为空的情况做了兼容,避免字符串格式化报错。
内容的提问来源于stack exchange,提问作者user23970789
相关产品推荐
相关产品推荐

