You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取XML文件中country节点的子节点文本值?

Python提取XML中country子节点的文本值

问题描述

给定如下XML内容:

1 2008 The house is on the hill 4 2011 59900

需要提取所有country节点下子节点的文本值(忽略无文本内容的节点),最终得到结果:[1, 2008, 'The house is on the hill', 4, 2011, 59900],如何用Python实现?

解决方案:使用标准库xml.etree.ElementTree

Python内置的xml.etree.ElementTree可以直接完成这个需求,无需额外安装依赖,具体代码如下:

import xml.etree.ElementTree as ET

# XML内容(如果是本地文件,可替换为ET.parse('your_file.xml').getroot())
xml_str = '''
<data>
    <country name="Liechtenstein">
        <rank>1</rank>
        <year>2008</year>
        <gdppc>The house is on the hill</gdppc>
        <neighbor name="Austria" direction="E"/>
        <neighbor name="Switzerland" direction="W"/>
    </country>
    <country name="Singapore">
        <rank>4</rank>
        <year>2011</year>
        <gdppc>59900</gdppc>
        <neighbor name="Malaysia" direction="N"/>
    </country>
</data>
'''

# 解析XML
root = ET.fromstring(xml_str)

result_list = []

# 遍历所有country节点
for country_node in root.findall('country'):
    # 遍历当前country的所有子节点
    for child_node in country_node:
        # 提取并清洗文本内容
        node_text = child_node.text.strip() if child_node.text else ''
        if node_text:
            # 尝试转换为整数,保留原字符串如果转换失败
            try:
                result_list.append(int(node_text))
            except ValueError:
                result_list.append(node_text)

print(result_list)
# 输出结果:[1, 2008, 'The house is on the hill', 4, 2011, 59900]

关键步骤说明

  • 解析XML:用ET.fromstring()处理字符串格式的XML,文件格式则用ET.parse()
  • 节点遍历:通过findall('country')定位所有目标父节点,再逐个遍历其子节点
  • 文本过滤:去除文本两端空白后,跳过无有效内容的节点(比如示例中的neighbor节点)
  • 类型转换:自动将数字文本转为整数,非数字文本保留原格式,匹配需求的结果格式

内容的提问来源于stack exchange,提问作者eliza nyambu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 02:30:57