You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy XPath:筛选不含span.status且子a含rent.html的article元素

筛选符合条件的class="car"的article元素

这里提供两个能满足需求的XPath表达式,分别从不同定位逻辑出发:

直接定位符合条件的article

//article[contains(@class, 'car') and not(.//span[contains(@class, 'status')]) and .//a[contains(@href, 'rent.html')]]

表达式拆解:

  • contains(@class, 'car'):匹配class属性包含car的article元素(避免严格匹配@class='car'导致多class场景失效)
  • not(.//span[contains(@class, 'status')]):排除自身及子层级中存在class为status的span元素的article
  • .//a[contains(@href, 'rent.html')]:确保article下任意层级存在href属性包含rent.html的a元素

从已定位的a元素反向查找对应的article

如果你已经定位到符合href条件的a元素,可以用这个表达式向上追溯父级article:

//a[contains(@href, 'rent.html')]/ancestor::article[contains(@class, 'car') and not(.//span[contains(@class, 'status')])]

匹配结果说明

结合你提供的示例HTML,前两个article.car会被选中:

  • 第三个article的a元素href是buy.html,不符合href条件
  • 第四个article包含span.status,被not()条件排除

内容的提问来源于stack exchange,提问作者Adam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 08:55:22