如何在Python中提取XML标签并保留其层级与顺序?
提取XML标签的层级路径列表
我来帮你搞定这个需求!用xml.etree.ElementTree实现带层级结构的标签路径提取,核心是通过递归遍历元素并跟踪路径,再根据你的需求过滤出想要的路径。
完整实现代码
import xml.etree.ElementTree as ET # 你的XML内容 xml_content = '''<Collection variable="value"> <Genre variable="value"> <Timestamp>2017-05-15T18:14:07-05:00</Timestamp> <Date>2016-12-31</Date> <Identifier> <id>123456789</id> <Name> <BusinessName>AB & co</BusinessName> </Name> </Identifier> </Genre> </Collection>''' # 解析XML得到根元素 root = ET.fromstring(xml_content) def get_tag_paths(element, current_path, paths): for child in element: # 生成当前子元素的完整层级路径 child_path = f"{current_path}/{child.tag}" # 按照你的需求过滤:保留根的直接子元素,以及所有终端子元素(没有子标签的元素) if current_path == root.tag or len(child) == 0: paths.append(child_path) # 递归处理当前子元素的子标签 get_tag_paths(child, child_path, paths) # 初始化路径列表 tag_paths = [] # 从根元素开始遍历 get_tag_paths(root, root.tag, tag_paths) # 输出结果 print(tag_paths)
代码说明
- 解析XML:用
ET.fromstring()直接解析XML字符串(如果是本地文件可以用ET.parse()),得到根元素Collection。 - 递归遍历函数:
get_tag_paths负责遍历每个元素:- 对每个子元素生成从根到它的完整路径(用
/分隔层级)。 - 按照你的期望输出过滤路径:只保留根的直接子元素(比如
Collection/Genre),以及所有没有子标签的终端元素(比如Timestamp、id等)。 - 递归处理子元素的子标签,确保遍历所有层级。
- 对每个子元素生成从根到它的完整路径(用
- 收集结果:初始化空列表,调用函数后得到最终的路径列表。
运行这段代码,输出正好是你想要的:
['Collection/Genre', 'Collection/Genre/Timestamp', 'Collection/Genre/Date', 'Collection/Genre/Identifier/id', 'Collection/Genre/Identifier/Name/BusinessName']
灵活调整
如果你需要其他规则的路径(比如保留所有层级的标签路径),只需要去掉if判断,直接把所有child_path加入列表即可,这样会得到包括Collection/Genre/Identifier、Collection/Genre/Identifier/Name在内的所有标签路径。
内容的提问来源于stack exchange,提问作者Okroshiashvili
相关产品推荐
相关产品推荐

