如何用Python 3.8.10将XML员工数据转为指定结构的Pandas DataFrame
问题:将XML员工数据转换为指定结构的Pandas DataFrame
我有存储在
employeedata.xml文件中的员工数据,样本数据如下:<?xml version="1.0" standalone="yes"?> <DocumentElement> <_x005B_dbo_x005D_._x005B_employeedata_x005D_> <RowID>11148</RowID> <ItemID>966109</ItemID> <Mappings>[]</Mappings> <Groups>93664,68349</Groups> <GroupKey>7003142</GroupKey> <ParentItemID>351908</ParentItemID> <JobID>30318</JobID> <Action>Employee</Action> <Employee_Name>John Travis</Employee_Name> <mail_id>John@altmail.com</mail_id> <...>...</...> <Action>Experience</Action> <...>...</...> <...>...</...> <...>...</...> </_x005B_dbo_x005D_._x005B_employeedata_x005D_> </DocumentElement>当前用Python 3.8.10读取文件的代码如下:
with open('employeedata.xml', 'r') as f: data = f.read()希望将这些XML数据存储到Pandas DataFrame中,将
<Action>标签的值存入Action列,对应的其他标签名存入Field列,预期输出如下:
Action Field Employee Employee_Name Employee mail_id Employee .... Employee .... Experience .... Experience ....
解决方案
实现思路
- 用Python标准库
xml.etree.ElementTree解析XML,无需额外安装依赖 - 遍历XML子节点,实时跟踪当前的
Action值 - 将每个非
Action节点的标签名与当前Action配对,存入列表后转换为DataFrame
代码实现
import xml.etree.ElementTree as ET import pandas as pd # 解析XML文件 tree = ET.parse('employeedata.xml') root = tree.getroot() # 定位到员工数据的父节点 employee_data_node = root.find('_x005B_dbo_x005D_._x005B_employeedata_x005D_') # 存储Action与Field的配对数据 result_list = [] current_action = None # 遍历所有子节点 for elem in employee_data_node: tag_name = elem.tag if tag_name == 'Action': # 更新当前生效的Action值 current_action = elem.text else: # 跳过XML中的占位节点(根据实际数据调整过滤规则) if tag_name != '...': result_list.append({ 'Action': current_action, 'Field': tag_name }) # 转换为DataFrame df = pd.DataFrame(result_list) print(df)
关键说明
- 代码精准定位到目标父节点
_x005B_dbo_x005D_._x005B_employeedata_x005D_,确保只处理员工数据节点 - 遇到
Action标签时更新当前关联的Action值,后续所有非Action标签自动归属于该Action - 内置了对
<...>占位节点的过滤,可根据实际XML结构调整过滤条件 - 最终生成的DataFrame完全匹配预期结构,可直接用于后续数据处理
内容的提问来源于stack exchange,提问作者python learner
相关产品推荐
相关产品推荐

