Scrapy爬取无class属性span中价格失败求助
Elektra手机页面Scrapy爬虫价格提取问题排查
问题描述
为网站https://www.elektra.mx/telefonia/celulares开发Scrapy爬虫时,可正常获取商品标题,但无法提取价格信息。目标价格所在<span>标签无class属性,且浏览器中可用的XPath在爬虫执行时返回空值。
相关代码
imgCss = response.xpath("(//img[contains(@class, 'vtex-product-summary-2-x-imageNormal')]/@src)[2]").get() title = response.xpath("(//article)[3]//span[contains(@class, 'vtex-product-summary-2-x-productBrand')]/text()").get() discount = response.xpath("(//article)[3]//span[contains(@class, 'currencyContainer--summary txt-price-responsive')]//text()").get() price = response.xpath("(//article)[3]//span[contains(@class, 'currencyContainer--summary t-heading-2-s')]//text()").get()
响应截图

排查与解决建议
- 动态渲染验证:Elektra这类电商站点常采用前端框架动态渲染价格数据,Scrapy默认获取的是初始静态HTML,可能未包含价格内容。可以通过打印
response.text确认价格是否存在,若不存在,需使用scrapy-splash或Playwright等工具渲染页面后再提取。 - 优化定位逻辑:依赖
(//article)[3]这种固定索引定位极不稳定,页面结构变动就会失效。建议先定位商品卡片容器(例如包含vtex-product-summary-2-x-container类的元素),再在容器内嵌套提取价格:# 示例:基于商品卡片容器的价格提取 product_cards = response.xpath("//div[contains(@class, 'vtex-product-summary-2-x-container')]") for card in product_cards: title = card.xpath(".//span[contains(@class, 'vtex-product-summary-2-x-productBrand')]/text()").get() # 针对无class的价格span,通过父容器或兄弟节点定位 price = card.xpath(".//div[contains(@class, 'priceContainer')]/span[not(@class)]/text()").get() - 无class标签定位技巧:若目标span无class,可通过上下文特征定位,比如:
- 找紧邻货币符号的span
- 通过父元素的class属性缩小范围,再选取对应位置的span
- 使用
text()内容匹配(如果价格前有固定文本提示)
内容的提问来源于stack exchange,提问作者Caliche Orozco
相关产品推荐
相关产品推荐

