如何将Pandas DataFrame转换为指定属性格式的嵌套XML?
解决方案
1. 单个DataFrame转属性式XML
先实现一个基础函数,把DataFrame转换成要求的<frameName>包裹、每行作为带属性子元素的XML片段:
import pandas as pd def df_to_attr_xml(df, tag_name): # 生成每行的XML元素 rows_xml = [] for _, row in df.iterrows(): # 拼接属性:列名="值",统一将值转为字符串并添加引号 attrs = ' '.join([f'{col}="{str(val)}"' for col, val in row.items()]) rows_xml.append(f' <{tag_name} {attrs} />') # 包裹外层标签 return f'<{tag_name}>\n' + '\n'.join(rows_xml) + f'\n</{tag_name}>' # 测试基础转换 df = pd.DataFrame({'category': ['A', 'B', 'C', 'D'], 'descr': ['blah', 'smthn', "yes", 'hello'], 'num1': [0,1,5,4], 'num2': [5,3,7,9]}) print(df_to_attr_xml(df, 'MyData'))
运行输出:
<MyData> <MyData category="A" descr="blah" num1="0" num2="5" /> <MyData category="B" descr="smthn" num1="1" num2="3" /> <MyData category="C" descr="yes" num1="5" num2="7" /> <MyData category="D" descr="hello" num1="4" num2="9" /> </MyData>
2. 生成嵌套XML结构
用嵌套字典定义整个XML的层级结构,再通过递归函数生成完整XML。字典规则:
tag: 当前节点的标签名attrs: 当前节点的属性(可选,字典格式)children: 子节点列表,每个子节点可以是另一个结构字典,或是DataFrame转换后的XML片段
递归生成函数:
def build_nested_xml(structure, indent=0): indent_str = ' ' * indent tag = structure['tag'] # 处理当前节点属性 attrs_str = '' if 'attrs' in structure: attrs_str = ' ' + ' '.join([f'{k}="{v}"' for k, v in structure['attrs'].items()]) # 处理子节点 children_xml = '' if 'children' in structure: for child in structure['children']: if isinstance(child, dict): # 递归处理子节点结构 children_xml += build_nested_xml(child, indent + 1) elif isinstance(child, str): # 给预生成的XML片段添加对应缩进 indented_child = '\n'.join([f'{indent_str} {line}' for line in child.split('\n')]) children_xml += f'\n{indented_child}' # 组装当前节点 if children_xml: return f'{indent_str}<{tag}{attrs_str}>{children_xml}\n{indent_str}</{tag}>' else: return f'{indent_str}<{tag}{attrs_str} />' # 准备嵌套结构所需的测试数据 some_df = pd.DataFrame({'col1': ['yes', 'no'], 'col2': ['hello', 'bye']}) some_frame_xml = df_to_attr_xml(some_df, 'someFrame') # 定义嵌套结构 nested_structure = { 'tag': 'highestCategory', 'attrs': {'fact1': '5', 'fact2': '8'}, 'children': [ { 'tag': 'lowerCategory', 'attrs': {'id': '69', 'details': 'abcd'}, 'children': [ some_frame_xml, df_to_attr_xml(df, 'MyData') ] } ] } # 生成完整嵌套XML full_xml = build_nested_xml(nested_structure) print(full_xml)
运行输出:
<highestCategory fact1="5" fact2="8"> <lowerCategory id="69" details="abcd"> <someFrame> <someFrame col1="yes" col2="hello" /> <someFrame col1="no" col2="bye" /> </someFrame> <MyData> <MyData category="A" descr="blah" num1="0" num2="5" /> <MyData category="B" descr="smthn" num1="1" num2="3" /> <MyData category="C" descr="yes" num1="5" num2="7" /> <MyData category="D" descr="hello" num1="4" num2="9" /> </MyData> </lowerCategory> </highestCategory>
补充说明
- 基础函数直接遍历DataFrame行拼接属性字符串,完全规避了
df.to_dict生成子节点的问题,精准匹配需求格式。 - 嵌套结构支持任意层级扩展,只需按字典规则定义结构即可。
- 若需处理特殊数据类型(如日期、布尔值),可在
df_to_attr_xml的属性拼接步骤中添加对应类型转换逻辑。
内容的提问来源于stack exchange,提问作者JPErwin
相关产品推荐
相关产品推荐

