如何用XPath/CSS定位器提取无唯一标识的HTML段落文本?
解决方法
一、定位单个特定段落
针对抓取特定标签(比如Hook:对应的段落)的场景,修正后的XPath和CSS选择器如下:
XPath 写法
通过子元素strong的文本内容定位父级p:
# 匹配包含指定文本的strong所在的p //div[contains(@class, 'entry-content')]/p[strong[contains(text(), 'Hook:')]] # 更精准的开头匹配(避免类似"Hooked:"的误匹配) //div[contains(@class, 'entry-content')]/p[strong[starts-with(text(), 'Hook:')]]
提取文本的示例代码:
# 获取段落完整文本 response.xpath("//div[contains(@class, 'entry-content')]/p[strong[starts-with(text(), 'Hook:')]]").get() # 仅提取strong标签后的内容(排除"Hook:") response.xpath("//div[contains(@class, 'entry-content')]/p[strong[starts-with(text(), 'Hook:')]]/text()").get().strip()
CSS 选择器写法
利用:has()伪类定位包含指定文本strong的p:
div.entry-content p:has(strong:contains('Hook:'))
提取文本的示例代码:
# 获取段落完整文本 response.css("div.entry-content p:has(strong:contains('Hook:'))").get() # 仅提取strong标签后的内容 response.css("div.entry-content p:has(strong:contains('Hook:'))::text").get().strip()
二、批量提取所有段落的键值对
如果需要一次性抓取所有p的标签(strong内文本)和对应内容,可批量处理:
XPath 实现
p_nodes = response.xpath("//div[contains(@class, 'entry-content')]/p[strong]") for node in p_nodes: label = node.xpath("./strong/text()").get().strip() content = node.xpath("./text()").get().strip() print(f"{label} {content}")
CSS 实现
p_nodes = response.css("div.entry-content p:has(strong)") for node in p_nodes: label = node.css("strong::text").get().strip() content = node.css("::text").get().strip() print(f"{label} {content}")
原XPath失败原因说明
//p[contains(@strong, 'Hook')]:@strong是引用元素的属性,但这里strong是子元素而非属性,语法错误。//p[contains(text(),'Inciting Event')]:text()仅获取p的直接文本节点,而"Inciting Event"在strong子元素内部,不在p的直接文本范围内,因此无法匹配。
内容的提问来源于stack exchange,提问作者mediacraftsman
相关产品推荐
相关产品推荐

