You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将指定格式的XML文件解析为DataFrame?

解决XML解析问题,输出指定CSV格式

我来帮你搞定这个XML解析的问题!先看看你现有代码的问题,再给出能达到预期效果的实现方案。

原代码的问题分析

  1. 错误的属性获取逻辑:你写的Age = attribute.attrib.get('attributes')完全找错了方向——<attribute>节点的属性是name和class,属性的实际值是节点的文本内容(比如<attribute name="age">25</attribute>里的25)。
  2. 未按记录分组处理:每个<attributes>节点对应一条完整的人员记录,你直接遍历所有<attribute>节点会把不同记录的属性混在一起,没法拼成完整的行。
  3. 未匹配字段与对应值:你没有根据name属性来对应age、carAvailability等指定字段,导致无法按要求的顺序输出内容。

正确的Python解析代码

用Python内置的xml.etree.ElementTree库就能轻松实现需求,代码逻辑清晰,完全匹配你的预期格式:

import xml.etree.ElementTree as ET

# 解析XML(如果是本地文件,可替换为ET.parse('your_xml_file.xml').getroot())
xml_content = '''
<person id="10000115">
  <attributes>
    <attribute name="age" class="java.lang.Integer" >25</attribute>
    <attribute name="carAvailability" class="java.lang.String" >never</attribute>
    <attribute name="censusId" class="java.lang.Integer" >449528</attribute>
    <attribute name="employment" class="java.lang.String" >yes</attribute>
    <attribute name="htsId" class="java.lang.String" >1009431</attribute>
    <attribute name="sex" class="java.lang.String" >m</attribute>
  </attributes>
  <attributes>
    <attribute name="age" class="java.lang.Integer" >55</attribute>
    <attribute name="carAvailability" class="java.lang.String" >never</attribute>
    <attribute name="censusId" class="java.lang.Integer" >450886</attribute>
    <attribute name="employment" class="java.lang.String" >yes</attribute>
    <attribute name="htsId" class="java.lang.String" >1023573</attribute>
    <attribute name="sex" class="java.lang.String" >m</attribute>
  </attributes>
</person>
'''
root = ET.fromstring(xml_content)

# 定义输出的字段顺序,和你期望的完全一致
output_fields = ['id', 'age', 'carAvailability', 'censusId', 'employment', 'htsId', 'sex']

# 打印表头
print(','.join(output_fields))

# 遍历每个person节点
for person in root.iter('person'):
    person_id = person.attrib['id']
    # 遍历每个attributes节点(每个节点对应一条独立记录)
    for attr_group in person.findall('attributes'):
        # 初始化记录字典,先存入person的id
        record = {'id': person_id}
        # 收集当前记录的所有属性值
        for attr in attr_group.findall('attribute'):
            attr_name = attr.attrib['name']
            attr_value = attr.text.strip()  # 去除文本前后的空格
            record[attr_name] = attr_value
        # 按指定字段顺序拼接成CSV行并打印
        csv_row = [record[field] for field in output_fields]
        print(','.join(csv_row))

代码说明

  1. XML解析方式:示例用ET.fromstring()解析字符串形式的XML,如果是本地XML文件,替换为tree = ET.parse('你的文件名.xml')后再root = tree.getroot()即可。
  2. 字段顺序控制:提前定义output_fields确保输出的列顺序和你要求的完全一致。
  3. 记录分组收集:每个<attributes>节点对应一条记录,先获取person的id,再逐个收集该记录的所有属性值。
  4. 属性值匹配:通过attribute节点的name属性,把对应文本值存入字典,最后按字段顺序提取输出。

输出结果

运行代码后会得到如下内容(和你的预期一致,仅第二条记录的id因原XML只有一个person节点暂用同一个id,若XML是多个person节点,代码无需修改即可自动适配):

id,age,carAvailability,censusId,employment,htsId,sex
10000115,25,never,449528,yes,1009431,m
10000115,55,never,450886,yes,1023573,m

内容的提问来源于stack exchange,提问作者nne

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 11:17:33