You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取无class属性span中价格失败求助

Elektra手机页面Scrapy爬虫价格提取问题排查

问题描述

为网站https://www.elektra.mx/telefonia/celulares开发Scrapy爬虫时,可正常获取商品标题,但无法提取价格信息。目标价格所在<span>标签无class属性,且浏览器中可用的XPath在爬虫执行时返回空值。

相关代码

imgCss = response.xpath("(//img[contains(@class, 'vtex-product-summary-2-x-imageNormal')]/@src)[2]").get()
title = response.xpath("(//article)[3]//span[contains(@class, 'vtex-product-summary-2-x-productBrand')]/text()").get()
discount = response.xpath("(//article)[3]//span[contains(@class, 'currencyContainer--summary txt-price-responsive')]//text()").get()
price = response.xpath("(//article)[3]//span[contains(@class, 'currencyContainer--summary t-heading-2-s')]//text()").get()

响应截图

商品价格标签截图

排查与解决建议

  • 动态渲染验证:Elektra这类电商站点常采用前端框架动态渲染价格数据,Scrapy默认获取的是初始静态HTML,可能未包含价格内容。可以通过打印response.text确认价格是否存在,若不存在,需使用scrapy-splash或Playwright等工具渲染页面后再提取。
  • 优化定位逻辑:依赖(//article)[3]这种固定索引定位极不稳定,页面结构变动就会失效。建议先定位商品卡片容器(例如包含vtex-product-summary-2-x-container类的元素),再在容器内嵌套提取价格:
    # 示例:基于商品卡片容器的价格提取
    product_cards = response.xpath("//div[contains(@class, 'vtex-product-summary-2-x-container')]")
    for card in product_cards:
        title = card.xpath(".//span[contains(@class, 'vtex-product-summary-2-x-productBrand')]/text()").get()
        # 针对无class的价格span,通过父容器或兄弟节点定位
        price = card.xpath(".//div[contains(@class, 'priceContainer')]/span[not(@class)]/text()").get()
    
  • 无class标签定位技巧:若目标span无class,可通过上下文特征定位,比如:
    • 找紧邻货币符号的span
    • 通过父元素的class属性缩小范围,再选取对应位置的span
    • 使用text()内容匹配(如果价格前有固定文本提示)

内容的提问来源于stack exchange,提问作者Caliche Orozco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 22:30:52