Python动态转换含数组与属性的嵌套XML至CSV方案问询
动态嵌套XML转CSV(Pandas实现)
XML示例
示例1:带属性的嵌套数组结构
<data> <country id="US"> <name>United States</name> <indicators> <indicator id="GDP"> <year>2022</year> <value>25.46</value> </indicator> <indicator id="CPI"> <year>2022</year> <value>8.0</value> </indicator> </indicators> </country> <country id="CN"> <name>China</name> <indicators> <indicator id="GDP"> <year>2022</year> <value>18.10</value> </indicator> <indicator id="POP"> <year>2022</year> <value>1412</value> </indicator> </indicators> </country> </data>
示例2:多类型子节点结构
<inventory> <product sku="P001"> <name>Laptop</name> <category>Electronics</category> <specs> <spec key="screen_size">15.6"</spec> <spec key="ram">16GB</spec> </specs> <prices> <price type="retail">$999</price> <price type="wholesale">$850</price> </prices> </product> <product sku="P002"> <name>Wireless Headphones</name> <category>Audio</category> <specs> <spec key="battery">30h</spec> </specs> <prices> <price type="retail">$199</price> </prices> </product> </inventory>
解决方案代码
import xml.etree.ElementTree as ET import pandas as pd def flatten_xml(node, parent_path="", data=None): if data is None: data = [] current_data = {} # 提取当前节点的所有属性,命名格式:父路径.节点名.属性名 for attr, val in node.attrib.items(): current_data[f"{parent_path}{node.tag}.{attr}"] = val # 提取当前节点的文本内容(忽略空白文本) if node.text and node.text.strip(): key = f"{parent_path}{node.tag}" if parent_path else node.tag current_data[key] = node.text.strip() # 按标签分组子节点,识别数组(同名子节点) child_groups = {} for child in node: child_groups.setdefault(child.tag, []).append(child) # 处理子节点 for tag, children in child_groups.items(): if len(children) > 1: # 数组节点:每个子元素生成独立记录,继承当前节点数据 for child in children: child_data = current_data.copy() flatten_xml(child, f"{parent_path}{node.tag}.", child_data) data.append(child_data) else: # 普通子节点:递归合并到当前数据 flatten_xml(children[0], f"{parent_path}{node.tag}.", current_data) # 无嵌套子节点时,将当前数据加入结果 if not child_groups and current_data: data.append(current_data) return data # 解析示例1并导出CSV tree = ET.parse("country_data.xml") root = tree.getroot() flat_data = flatten_xml(root) pd.DataFrame(flat_data).to_csv("country_data.csv", index=False) # 解析示例2并导出CSV tree = ET.parse("inventory.xml") root = tree.getroot() flat_data = flatten_xml(root) pd.DataFrame(flat_data).to_csv("inventory.csv", index=False)
关键特性说明
- 动态识别结构:无需硬编码标签名,自动适配任意嵌套XML
- 属性完整保留:节点属性会以
父节点.当前节点.属性名的格式作为CSV列名(如data.country.id) - 数组自动拆分:同名子节点被视为数组,每个元素生成独立行
- 多层嵌套扁平化:所有嵌套节点都会展开为一维键值对,缺失字段自动填充
NaN
内容的提问来源于stack exchange,提问作者Eja
相关产品推荐
相关产品推荐

