如何使用Python提取XML文件中country节点的子节点文本值?
Python提取XML中country子节点的文本值
问题描述
给定如下XML内容:
1 2008 The house is on the hill 4 2011 59900
需要提取所有country节点下子节点的文本值(忽略无文本内容的节点),最终得到结果:[1, 2008, 'The house is on the hill', 4, 2011, 59900],如何用Python实现?
解决方案:使用标准库xml.etree.ElementTree
Python内置的xml.etree.ElementTree可以直接完成这个需求,无需额外安装依赖,具体代码如下:
import xml.etree.ElementTree as ET # XML内容(如果是本地文件,可替换为ET.parse('your_file.xml').getroot()) xml_str = ''' <data> <country name="Liechtenstein"> <rank>1</rank> <year>2008</year> <gdppc>The house is on the hill</gdppc> <neighbor name="Austria" direction="E"/> <neighbor name="Switzerland" direction="W"/> </country> <country name="Singapore"> <rank>4</rank> <year>2011</year> <gdppc>59900</gdppc> <neighbor name="Malaysia" direction="N"/> </country> </data> ''' # 解析XML root = ET.fromstring(xml_str) result_list = [] # 遍历所有country节点 for country_node in root.findall('country'): # 遍历当前country的所有子节点 for child_node in country_node: # 提取并清洗文本内容 node_text = child_node.text.strip() if child_node.text else '' if node_text: # 尝试转换为整数,保留原字符串如果转换失败 try: result_list.append(int(node_text)) except ValueError: result_list.append(node_text) print(result_list) # 输出结果:[1, 2008, 'The house is on the hill', 4, 2011, 59900]
关键步骤说明
- 解析XML:用
ET.fromstring()处理字符串格式的XML,文件格式则用ET.parse() - 节点遍历:通过
findall('country')定位所有目标父节点,再逐个遍历其子节点 - 文本过滤:去除文本两端空白后,跳过无有效内容的节点(比如示例中的
neighbor节点) - 类型转换:自动将数字文本转为整数,非数字文本保留原格式,匹配需求的结果格式
内容的提问来源于stack exchange,提问作者eliza nyambu
相关产品推荐
相关产品推荐

