You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame转换为指定属性格式的嵌套XML?

解决方案

1. 单个DataFrame转属性式XML

先实现一个基础函数,把DataFrame转换成要求的<frameName>包裹、每行作为带属性子元素的XML片段:

import pandas as pd

def df_to_attr_xml(df, tag_name):
    # 生成每行的XML元素
    rows_xml = []
    for _, row in df.iterrows():
        # 拼接属性:列名="值",统一将值转为字符串并添加引号
        attrs = ' '.join([f'{col}="{str(val)}"' for col, val in row.items()])
        rows_xml.append(f'    <{tag_name} {attrs} />')
    # 包裹外层标签
    return f'<{tag_name}>\n' + '\n'.join(rows_xml) + f'\n</{tag_name}>'

# 测试基础转换
df = pd.DataFrame({'category': ['A', 'B', 'C', 'D'],
                   'descr': ['blah', 'smthn', "yes", 'hello'],
                   'num1': [0,1,5,4],
                   'num2': [5,3,7,9]})

print(df_to_attr_xml(df, 'MyData'))

运行输出:

<MyData>
    <MyData category="A" descr="blah" num1="0" num2="5" />
    <MyData category="B" descr="smthn" num1="1" num2="3" />
    <MyData category="C" descr="yes" num1="5" num2="7" />
    <MyData category="D" descr="hello" num1="4" num2="9" />
</MyData>

2. 生成嵌套XML结构

用嵌套字典定义整个XML的层级结构,再通过递归函数生成完整XML。字典规则:

  • tag: 当前节点的标签名
  • attrs: 当前节点的属性(可选,字典格式)
  • children: 子节点列表,每个子节点可以是另一个结构字典,或是DataFrame转换后的XML片段

递归生成函数:

def build_nested_xml(structure, indent=0):
    indent_str = '    ' * indent
    tag = structure['tag']
    # 处理当前节点属性
    attrs_str = ''
    if 'attrs' in structure:
        attrs_str = ' ' + ' '.join([f'{k}="{v}"' for k, v in structure['attrs'].items()])
    # 处理子节点
    children_xml = ''
    if 'children' in structure:
        for child in structure['children']:
            if isinstance(child, dict):
                # 递归处理子节点结构
                children_xml += build_nested_xml(child, indent + 1)
            elif isinstance(child, str):
                # 给预生成的XML片段添加对应缩进
                indented_child = '\n'.join([f'{indent_str}    {line}' for line in child.split('\n')])
                children_xml += f'\n{indented_child}'
    # 组装当前节点
    if children_xml:
        return f'{indent_str}<{tag}{attrs_str}>{children_xml}\n{indent_str}</{tag}>'
    else:
        return f'{indent_str}<{tag}{attrs_str} />'

# 准备嵌套结构所需的测试数据
some_df = pd.DataFrame({'col1': ['yes', 'no'], 'col2': ['hello', 'bye']})
some_frame_xml = df_to_attr_xml(some_df, 'someFrame')

# 定义嵌套结构
nested_structure = {
    'tag': 'highestCategory',
    'attrs': {'fact1': '5', 'fact2': '8'},
    'children': [
        {
            'tag': 'lowerCategory',
            'attrs': {'id': '69', 'details': 'abcd'},
            'children': [
                some_frame_xml,
                df_to_attr_xml(df, 'MyData')
            ]
        }
    ]
}

# 生成完整嵌套XML
full_xml = build_nested_xml(nested_structure)
print(full_xml)

运行输出:

<highestCategory fact1="5" fact2="8">
    <lowerCategory id="69" details="abcd">
        <someFrame>
            <someFrame col1="yes" col2="hello" />
            <someFrame col1="no" col2="bye" />
        </someFrame>
        <MyData>
            <MyData category="A" descr="blah" num1="0" num2="5" />
            <MyData category="B" descr="smthn" num1="1" num2="3" />
            <MyData category="C" descr="yes" num1="5" num2="7" />
            <MyData category="D" descr="hello" num1="4" num2="9" />
        </MyData>
    </lowerCategory>
</highestCategory>

补充说明

  • 基础函数直接遍历DataFrame行拼接属性字符串,完全规避了df.to_dict生成子节点的问题,精准匹配需求格式。
  • 嵌套结构支持任意层级扩展,只需按字典规则定义结构即可。
  • 若需处理特殊数据类型(如日期、布尔值),可在df_to_attr_xml的属性拼接步骤中添加对应类型转换逻辑。

内容的提问来源于stack exchange,提问作者JPErwin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 04:53:18