You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy CSS选择器提取首个指定类文本失败求助

解决Scrapy提取第一个同名类文本的问题

以下是几种可行的解决方案,你可以逐个尝试:

1. 直接定位第一个目标元素(CSS选择器)

用:first-child伪类直接选中第一个.s-recipe-header__info-item元素,简化选择器层级:

recipe_item["preparation_time"] = response.css(".s-recipe-header__info-item:first-child::text").get()

如果文本嵌套在该元素的子标签里(比如<span>),需要获取所有后代文本:

recipe_item["preparation_time"] = response.css(".s-recipe-header__info-item:first-child *::text").get()

2. 修正原选择器的层级关系

原选择器用了>(直接子元素),如果页面DOM结构里,.s-recipe-header__info-items不是.s-recipe-header__info的直接子元素,或者.s-recipe-header__info-item不是前者的直接子元素,就会匹配失败。换成空格(后代选择器)试试:

recipe_item["preparation_time"] = response.css(".s-recipe-header__info .s-recipe-header__info-items .s-recipe-header__info-item:first-child::text").get()

3. 使用XPath选择器(更灵活)

XPath可以通过索引直接定位第一个匹配的元素,有时候比CSS更直观:

# 匹配第一个带目标类的元素,提取其文本
recipe_item["preparation_time"] = response.xpath("(//*[contains(@class, 's-recipe-header__info-item')])[1]/text()").get()
# 如果文本在子元素里,改成:
recipe_item["preparation_time"] = response.xpath("(//*[contains(@class, 's-recipe-header__info-item')])[1]//text()").get()

调试小技巧

建议用Scrapy Shell测试选择器,输入命令:

scrapy shell 你的目标页面URL

然后在Shell里逐个测试response.css(...)或response.xpath(...),看返回的结果是什么,能快速定位是选择器写错了,还是文本位置不对。

内容的提问来源于stack exchange,提问作者Elis_the_Fox

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 03:35:22