You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy使用CSS选择器无法获取目标文本内容的问题求助

Scrapy使用CSS选择器无法获取目标文本内容的问题求助

看起来你的问题出在CSS选择器写得太“死板”了——你用到了elementor-44463、elementor-element-49bd7747这类ID,这些都是Elementor页面编辑器动态生成的唯一标识,不仅页面更新后可能会变化,而且只能精准匹配那一个特定元素,自然拿不到所有你需要的培训标签内容。

我们来看目标HTML的结构:所有你要抓取的文本,都嵌套在.elementor-widget-heading这个组件的.elementor-heading-title.elementor-size-default标签里。我们可以用这个通用的结构来写选择器,完全不用依赖那些不稳定的动态ID。

给你修改后的完整代码:

import scrapy

class monavenirpro(scrapy.Spider):
    name = "monavenirpro"
    start_urls = ["https://monavenirpro-hdf.fr/formations/"]

    def parse(self, response):
        # 匹配所有符合结构的目标标题文本
        libelles = response.css(".elementor-widget-heading .elementor-heading-title.elementor-size-default::text").extract()
        yield {"listlibelles": libelles}

运行这段代码后,就能抓取到所有你需要的内容,比如Accompagnement、Validation des Acquis de l'Expérience - Livret 2这类文本了。

如果你后续需要区分层级(比如把分类标题和具体培训项分开),还可以用更精细的选择器:

  • 只抓取具体的培训项(每个分类下的第二个标题):
    training_items = response.css(".elementor-widget-heading + .elementor-widget-heading .elementor-heading-title.elementor-size-default::text").extract()
    
  • 只抓取分类标题(每个组的第一个标题):
    categories = response.css(".elementor-widget-heading:nth-child(odd) .elementor-heading-title.elementor-size-default::text").extract()
    

备注:内容来源于stack exchange,提问作者Christelle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 10:44:33