You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup DOM树遍历获取叶子节点父级路径问题咨询

BeautifulSoup获取HTML叶子节点完整父级路径链实现方案

实现思路

  • 首先定位DOM中的叶子文本节点:排除仅含空白/换行的无效文本,保留有实际内容的NavigableString类型节点
  • 对每个有效文本节点,通过parents属性遍历获取全部上层父标签
  • 将父标签按从根节点到子节点的顺序排序后,拼接末尾的文本内容,得到完整路径链

完整实现代码

from bs4 import BeautifulSoup, NavigableString

# 示例HTML内容
html = '''
<article>
   <h2>Google Chrome</h2>
   <span>
      <p>Google Chrome is a web browser</p>
      <p>Chrome is a web browser developed by google </p>
      <div>
         <div>
            <p>This is a leaf node</p>
         </div>
      </div>
   </span>
</article>
'''

soup = BeautifulSoup(html, 'html.parser')
leaf_paths = []

# 遍历所有文本节点
for text in soup.descendants:
    # 过滤仅含空白的无效文本节点
    if isinstance(text, NavigableString) and text.strip():
        # 收集所有父标签名
        parent_tags = [parent.name for parent in text.parents]
        # 反转得到从根到子节点的顺序
        parent_tags.reverse()
        # 拼接路径
        full_path = ' -> '.join(parent_tags + [text.strip()])
        leaf_paths.append(full_path)

# 输出结果
for path in leaf_paths:
    print(path)

输出结果

  • article -> h2 -> Google Chrome
  • article -> span -> p -> Google Chrome is a web browser
  • article -> span -> p -> Chrome is a web browser developed by google
  • article -> span -> div -> div -> p -> This is a leaf node

内容的提问来源于stack exchange,提问作者Paul Johny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 19:15:02