如何用Python将指定格式的XML文件解析为DataFrame?
解决XML解析问题,输出指定CSV格式
我来帮你搞定这个XML解析的问题!先看看你现有代码的问题,再给出能达到预期效果的实现方案。
原代码的问题分析
- 错误的属性获取逻辑:你写的
Age = attribute.attrib.get('attributes')完全找错了方向——<attribute>节点的属性是name和class,属性的实际值是节点的文本内容(比如<attribute name="age">25</attribute>里的25)。 - 未按记录分组处理:每个
<attributes>节点对应一条完整的人员记录,你直接遍历所有<attribute>节点会把不同记录的属性混在一起,没法拼成完整的行。 - 未匹配字段与对应值:你没有根据
name属性来对应age、carAvailability等指定字段,导致无法按要求的顺序输出内容。
正确的Python解析代码
用Python内置的xml.etree.ElementTree库就能轻松实现需求,代码逻辑清晰,完全匹配你的预期格式:
import xml.etree.ElementTree as ET # 解析XML(如果是本地文件,可替换为ET.parse('your_xml_file.xml').getroot()) xml_content = ''' <person id="10000115"> <attributes> <attribute name="age" class="java.lang.Integer" >25</attribute> <attribute name="carAvailability" class="java.lang.String" >never</attribute> <attribute name="censusId" class="java.lang.Integer" >449528</attribute> <attribute name="employment" class="java.lang.String" >yes</attribute> <attribute name="htsId" class="java.lang.String" >1009431</attribute> <attribute name="sex" class="java.lang.String" >m</attribute> </attributes> <attributes> <attribute name="age" class="java.lang.Integer" >55</attribute> <attribute name="carAvailability" class="java.lang.String" >never</attribute> <attribute name="censusId" class="java.lang.Integer" >450886</attribute> <attribute name="employment" class="java.lang.String" >yes</attribute> <attribute name="htsId" class="java.lang.String" >1023573</attribute> <attribute name="sex" class="java.lang.String" >m</attribute> </attributes> </person> ''' root = ET.fromstring(xml_content) # 定义输出的字段顺序,和你期望的完全一致 output_fields = ['id', 'age', 'carAvailability', 'censusId', 'employment', 'htsId', 'sex'] # 打印表头 print(','.join(output_fields)) # 遍历每个person节点 for person in root.iter('person'): person_id = person.attrib['id'] # 遍历每个attributes节点(每个节点对应一条独立记录) for attr_group in person.findall('attributes'): # 初始化记录字典,先存入person的id record = {'id': person_id} # 收集当前记录的所有属性值 for attr in attr_group.findall('attribute'): attr_name = attr.attrib['name'] attr_value = attr.text.strip() # 去除文本前后的空格 record[attr_name] = attr_value # 按指定字段顺序拼接成CSV行并打印 csv_row = [record[field] for field in output_fields] print(','.join(csv_row))
代码说明
- XML解析方式:示例用
ET.fromstring()解析字符串形式的XML,如果是本地XML文件,替换为tree = ET.parse('你的文件名.xml')后再root = tree.getroot()即可。 - 字段顺序控制:提前定义
output_fields确保输出的列顺序和你要求的完全一致。 - 记录分组收集:每个
<attributes>节点对应一条记录,先获取person的id,再逐个收集该记录的所有属性值。 - 属性值匹配:通过
attribute节点的name属性,把对应文本值存入字典,最后按字段顺序提取输出。
输出结果
运行代码后会得到如下内容(和你的预期一致,仅第二条记录的id因原XML只有一个person节点暂用同一个id,若XML是多个person节点,代码无需修改即可自动适配):
id,age,carAvailability,censusId,employment,htsId,sex 10000115,25,never,449528,yes,1009431,m 10000115,55,never,450886,yes,1023573,m
内容的提问来源于stack exchange,提问作者nne
相关产品推荐
相关产品推荐

