Python BeautifulSoup DOM树遍历获取叶子节点父级路径问题咨询
BeautifulSoup获取HTML叶子节点完整父级路径链实现方案
实现思路
- 首先定位DOM中的叶子文本节点:排除仅含空白/换行的无效文本,保留有实际内容的
NavigableString类型节点 - 对每个有效文本节点,通过
parents属性遍历获取全部上层父标签 - 将父标签按从根节点到子节点的顺序排序后,拼接末尾的文本内容,得到完整路径链
完整实现代码
from bs4 import BeautifulSoup, NavigableString # 示例HTML内容 html = ''' <article> <h2>Google Chrome</h2> <span> <p>Google Chrome is a web browser</p> <p>Chrome is a web browser developed by google </p> <div> <div> <p>This is a leaf node</p> </div> </div> </span> </article> ''' soup = BeautifulSoup(html, 'html.parser') leaf_paths = [] # 遍历所有文本节点 for text in soup.descendants: # 过滤仅含空白的无效文本节点 if isinstance(text, NavigableString) and text.strip(): # 收集所有父标签名 parent_tags = [parent.name for parent in text.parents] # 反转得到从根到子节点的顺序 parent_tags.reverse() # 拼接路径 full_path = ' -> '.join(parent_tags + [text.strip()]) leaf_paths.append(full_path) # 输出结果 for path in leaf_paths: print(path)
输出结果
- article -> h2 -> Google Chrome
- article -> span -> p -> Google Chrome is a web browser
- article -> span -> p -> Chrome is a web browser developed by google
- article -> span -> div -> div -> p -> This is a leaf node
内容的提问来源于stack exchange,提问作者Paul Johny
相关产品推荐
相关产品推荐

