如何用XPath选择span内第一个<hr>标签之后的所有内容?
提取span内第一个
<hr>之后的所有内容(保留格式) 需要从id为selectorID的<span>元素中,提取第一个<hr>标签之后的所有内容——包括嵌套的列表、链接、后续<hr>及对应文本,且要保留原有的HTML结构和格式,不能仅提取纯文本。
示例HTML结构:
<span id="selectorID"> <b>Header Text</b> Some more header text. <hr> Body text that I want starts here, it may also include <a href="www.google.com">links</a>, <b>bolded text</b>, and even... <ul> <li>Lists!</li> <li>With a bunch of items.</li> <li>I want these too.</li> </ul> Then after all of that, it may also include <hr> Another HR, <b>but I want this text too that comes after this.</b> As long as it's after the first hr. </span>
之前尝试的XPath只能提取部分文本,无法获取嵌套标签内的内容:
//span[contains(@id,'selectorID')]/descendant-or-self::*/text()[count(preceding-sibling::hr)>0]
正确实现方式
适用于XPath 1.0(多数环境兼容)
使用以下XPath表达式,可选中所有目标节点(包括嵌套元素和文本):
//span[@id='selectorID']/node()[count(preceding-sibling::hr) > 0 or count(ancestor::*[preceding-sibling::hr]) > 0]
逻辑说明:
- 选中span下自身是第一个
<hr>后续兄弟的节点 - 同时选中祖先节点是第一个
<hr>后续兄弟的节点(覆盖嵌套在列表、链接里的内容)
适用于XPath 2.0+(更简洁)
如果环境支持XPath 2.0及以上,可使用更直观的写法:
//span[@id='selectorID']/(hr/following-sibling::node())//node()
逻辑说明:
- 先定位到span下第一个
<hr>的所有后续兄弟节点 - 再选中这些节点及其所有子节点,完整覆盖所有目标内容
注意事项
获取节点集合后,保留它们的层级关系进行处理,就能维持原有的换行、格式和HTML结构,无需依赖text()提取纯文本。
内容的提问来源于stack exchange,提问作者Twiggies
相关产品推荐
相关产品推荐

