You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用XPATH获取两个h3标签之间的所有元素

提取h3标签间内容的XPath方案

单个区间提取(第一个h3到第二个h3)

直接用XPath就能精准定位两个h3之间的所有同级元素,推荐两种实用写法:

  1. 基于前置h3数量筛选
//article/h3[1]/following-sibling::*[count(preceding-sibling::h3) = 1]
  • 逻辑:先定位第一个h3,再选取它之后的所有同级元素,最后筛选出前面仅有1个h3的元素——这些元素恰好处于第一个和第二个h3之间。
  1. 基于排除后续h3的前置元素
//article/h3[1]/following-sibling::*[not(preceding-sibling::h3[2])]
  • 逻辑:选取第一个h3之后的同级元素,排除那些位于第二个h3之后的元素,剩余部分就是两个h3之间的内容。

批量处理所有h3对应内容

如果要自动遍历所有h3,提取每个副标题和对应的后续内容,以Python的lxml库为例,代码示例如下:

from lxml import etree

# 假设html_content是你的文章HTML源码
tree = etree.HTML(html_content)
h3_nodes = tree.xpath('//article/h3')

for index, h3 in enumerate(h3_nodes):
    # 提取副标题文本
    subtitle = h3.xpath('string(.)').strip()
    # 提取对应内容
    if index < len(h3_nodes) - 1:
        # 非最后一个h3:取当前h3到下一个h3之间的元素
        content_nodes = h3.xpath(f'following-sibling::*[preceding-sibling::h3[{index+1}] and not(preceding-sibling::h3[{index+2}])]')
    else:
        # 最后一个h3:取之后所有同级元素
        content_nodes = h3.xpath('following-sibling::*')
    # 把元素转为纯文本
    content = ' '.join([etree.tostring(node, encoding='unicode', method='text').strip() for node in content_nodes])
    # 输出结果
    print(f"副标题:{subtitle}\n内容:{content}\n")

内容的提问来源于stack exchange,提问作者Prinple Vrlo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 13:15:51