You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用XPath/CSS定位器提取无唯一标识的HTML段落文本?

解决方法

一、定位单个特定段落

针对抓取特定标签(比如Hook:对应的段落)的场景,修正后的XPath和CSS选择器如下:

XPath 写法

通过子元素strong的文本内容定位父级p:

# 匹配包含指定文本的strong所在的p
//div[contains(@class, 'entry-content')]/p[strong[contains(text(), 'Hook:')]]

# 更精准的开头匹配(避免类似"Hooked:"的误匹配)
//div[contains(@class, 'entry-content')]/p[strong[starts-with(text(), 'Hook:')]]

提取文本的示例代码:

# 获取段落完整文本
response.xpath("//div[contains(@class, 'entry-content')]/p[strong[starts-with(text(), 'Hook:')]]").get()

# 仅提取strong标签后的内容(排除"Hook:")
response.xpath("//div[contains(@class, 'entry-content')]/p[strong[starts-with(text(), 'Hook:')]]/text()").get().strip()

CSS 选择器写法

利用:has()伪类定位包含指定文本strong的p:

div.entry-content p:has(strong:contains('Hook:'))

提取文本的示例代码:

# 获取段落完整文本
response.css("div.entry-content p:has(strong:contains('Hook:'))").get()

# 仅提取strong标签后的内容
response.css("div.entry-content p:has(strong:contains('Hook:'))::text").get().strip()

二、批量提取所有段落的键值对

如果需要一次性抓取所有p的标签(strong内文本)和对应内容,可批量处理:

XPath 实现

p_nodes = response.xpath("//div[contains(@class, 'entry-content')]/p[strong]")
for node in p_nodes:
    label = node.xpath("./strong/text()").get().strip()
    content = node.xpath("./text()").get().strip()
    print(f"{label} {content}")

CSS 实现

p_nodes = response.css("div.entry-content p:has(strong)")
for node in p_nodes:
    label = node.css("strong::text").get().strip()
    content = node.css("::text").get().strip()
    print(f"{label} {content}")

原XPath失败原因说明

  1. //p[contains(@strong, 'Hook')]:@strong是引用元素的属性,但这里strong是子元素而非属性,语法错误。
  2. //p[contains(text(),'Inciting Event')]:text()仅获取p的直接文本节点,而"Inciting Event"在strong子元素内部,不在p的直接文本范围内,因此无法匹配。

内容的提问来源于stack exchange,提问作者mediacraftsman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 16:55:42