Python ElementTree 筛选获取带指定属性的XML元素方法
提取XML中带指定属性元素的实现方案
直接通过attrib['population']取值时,若元素无该属性会抛出KeyError导致代码中断,可通过以下两种方式实现精准筛选,仅保留带population属性的目标元素:
方法1:XPath直接匹配(推荐)
ElementTree的findall方法原生支持XPath属性存在匹配规则,使用.//*[@population]语法可以直接定位全文档下所有携带population属性的节点,无需额外判空,代码最简洁:
from xml.etree import ElementTree as ET xml = '''<?xml version="1.0"?> <data> <country name="Liechtenstein"> <rank>1</rank> <year>2008</year> <gdppc>141100</gdppc> <neighbor name="Austria" direction="E"/> <neighbor name="Switzerland" direction="W"/> </country> <country name="Singapore"> <rank>4</rank> <year>2011</year> <gdppc>59900</gdppc> <neighbor name="Malaysia" direction="N"/> </country> <country name="Panama"> <rank>68</rank> <year>2011</year> <gdppc>13600</gdppc> <neighbor name="Costa Rica" direction="W"/> <neighbor name="Colombia" direction="E" population="500"/> </country> </data>''' root = ET.fromstring(xml) # 直接筛选所有带population属性的元素 for target in root.findall('.//*[@population]'): population = target.attrib['population'] print(f"标签:{target.tag},population值:{population},相邻方向:{target.attrib.get('direction')}")
如果明确
population属性只出现在特定标签(比如示例中的neighbor标签),可以将XPath规则收窄为.//neighbor[@population],匹配效率更高,也不会误匹配其他标签的同名属性。如果目标属性是挂载在country标签上,把规则改成.//country[@population]即可。
方法2:遍历节点后判断属性存在性
如果需要叠加其他自定义筛选逻辑,可以在遍历节点时先判断属性是否存在,再执行后续取值操作:
for country in root.findall('.//country'): for neighbor in country.findall('neighbor'): # 先判断属性是否存在,避免KeyError if 'population' in neighbor.attrib: neighbor_name = neighbor.attrib['name'] population = neighbor.attrib['population'] print(f"邻国{neighbor_name}的population值为{population}")
两种方案都会自动跳过无population属性的元素,不会触发属性不存在的报错,最终输出的全部是你需要的目标内容。
内容的提问来源于stack exchange,提问作者GrizzlyMcGee
相关产品推荐
相关产品推荐

