You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求通用Python脚本/教程:将多份不同字段嵌套XML转为DataFrame

通用Python脚本:复杂嵌套XML转DataFrame

核心思路

  1. 用xmltodict将XML解析为嵌套字典/列表结构,比原生xml.etree.ElementTree更易处理嵌套层级
  2. 递归扁平化嵌套结构,将多层嵌套的字段转为父节点.子节点格式的扁平字段
  3. 自动识别DataFrame中值为列表的字段,执行explode操作展开嵌套数组
  4. 最终将扁平数据转换为标准DataFrame,适配任意结构的嵌套XML

通用实现脚本

import xmltodict
import pandas as pd

def flatten_dict(nested_dict, parent_key='', sep='.'):
    """递归扁平化嵌套字典"""
    items = []
    for k, v in nested_dict.items():
        new_key = f"{parent_key}{sep}{k}" if parent_key else k
        if isinstance(v, dict):
            items.extend(flatten_dict(v, new_key, sep=sep).items())
        elif isinstance(v, list):
            # 列表类型暂存,后续统一处理explode
            items.append((new_key, v))
        else:
            items.append((new_key, v))
    return dict(items)

def xml_to_df(xml_content, sep='.'):
    """将嵌套XML转换为带explode处理的DataFrame"""
    # 解析XML为字典
    xml_dict = xmltodict.parse(xml_content)
    
    # 扁平化字典(根节点通常是XML的顶层标签,先取出)
    root_key = next(iter(xml_dict.keys()))
    flat_data = flatten_dict(xml_dict[root_key], sep=sep)
    
    # 转换为初始DataFrame
    df = pd.DataFrame([flat_data])
    
    # 自动识别并explode所有列表类型的字段
    list_columns = [col for col in df.columns if isinstance(df[col].iloc[0], list)]
    for col in list_columns:
        df = df.explode(col, ignore_index=True)
        # 若explode后的字段仍是嵌套结构,再次扁平化
        if isinstance(df[col].iloc[0], dict):
            exploded_df = df[col].apply(pd.Series).add_prefix(f"{col}{sep}")
            df = pd.concat([df.drop(col, axis=1), exploded_df], axis=1)
    
    return df

关键逻辑说明

  • 递归扁平化:flatten_dict函数遍历所有嵌套层级,将子节点字段名拼接成父节点.子节点格式,比如<user><name>Alice</name></user>会转为user.name: Alice
  • 自动explode处理:遍历DataFrame列,判断字段值是否为列表,对每个列表字段执行explode展开;如果展开后的元素仍是字典,会进一步将字典转为多列,彻底扁平化嵌套结构
  • 适配任意XML结构:无需针对不同XML字段修改代码,脚本会自动识别所有嵌套层级和数组字段

使用示例

假设你有如下嵌套XML:

<orders>
    <order>
        <order_id>1001</order_id>
        <customer>
            <name>John Doe</name>
            <email>john@example.com</email>
        </customer>
        <items>
            <item>
                <product>Apple</product>
                <quantity>2</quantity>
            </item>
            <item>
                <product>Banana</product>
                <quantity>5</quantity>
            </item>
        </items>
    </order>
</orders>

调用脚本处理:

# 读取XML文件(或直接传入XML字符串)
with open('orders.xml', 'r') as f:
    xml_content = f.read()

df = xml_to_df(xml_content)
print(df)

输出的DataFrame会展开items列表,生成2行数据,同时扁平化customer和item的嵌套字段:

order_id customer.name customer.email items.product items.quantity
0     1001       John Doe  john@example.com         Apple              2
1     1001       John Doe  john@example.com        Banana              5

注意事项

  • 确保已安装依赖库:pip install xmltodict pandas
  • 若XML中有重复的同级标签(如多个<order>),脚本会自动处理为列表并explode展开
  • 可修改sep参数自定义嵌套字段的分隔符(比如用_代替.)

内容的提问来源于stack exchange,提问作者DAMINI RANA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 23:54:58