Python如何提取XML中各类型标签对应的description汇总列表
实现思路
- 首先处理带命名空间的标签,通过
split('}')[-1]提取type1、type2这类短名称 - 用字典存储每个类型对应的描述列表,避免重复统计
- 遍历每个类型的所有节点,检索其下的
description子节点内容,按要求格式化输出即可
完整实现代码
import xml.etree.ElementTree as ET # 解析XML文件 test = ET.parse("...\\index.xml") root = test.getroot() # 获取所有去重的带命名空间标签(复用你原有逻辑) type_list = [] for elem in test.iter(): type_list.append(elem.tag) type_list = list(set(type_list)) result_dict = {} for full_tag in type_list: # 提取不带命名空间的短类型名 short_type = full_tag.split("}")[-1] # 过滤非type开头的无关标签(如batch、description等) if not short_type.startswith("type"): continue result_dict[short_type] = [] # 遍历当前类型的所有节点 for type_node in root.iter(full_tag): # 查找description子节点,如果description也带命名空间,此处替换为对应全名即可 # 例:如果description的tag是{blabla.com}description,就写为 type_node.find("{blabla.com}description") desc_node = type_node.find("description") if desc_node is not None and desc_node.text: result_dict[short_type].append(desc_node.text.strip()) # 按要求格式输出结果 for type_name, desc_list in result_dict.items(): print(f"{type_name}: {', '.join(desc_list)}")
注意事项
- 如果
description标签也带命名空间,可以先执行print([elem.tag for elem in test.iter() if 'description' in elem.tag])查看其完整标签格式,替换find方法的入参即可 - 如果需要只提取指定的type类型,可在过滤短名称的步骤中增加自定义判断条件
内容的提问来源于stack exchange,提问作者Conquering
相关产品推荐
相关产品推荐

