You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python xml.dom.minidom访问XML子子节点并导出数据到表格

XML数据提取及电子表格导出解决方案

问题背景

给定如下XML结构,需提取数据并导出为指定格式的电子表格,但原有代码无法访问grandchildren标签内容,需解决该问题。

源XML数据

<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<root testAttr="testValue">
    <family name="Hardwood">
        <child name="Jack">First</child>
        <child name="Rose">Second</child>
        <child name="Blue Ivy">Third
                <grandchildren>
                    <data>One</data>
                    <data>Two</data>
                    <unique>Twins</unique>
                </grandchildren>
            </child>
        <child name="Jane">Fourth</child>
    </family>
    <family name="Downie">
        <child name="Bill">First</child>
        <child name="Rosie">Second</child>
        <child name="Edward">
            Third
            </child>
        <child name="Jane">Fourth</child>
    </family>
</root>

原有尝试代码

import xml.dom.minidom
doc=xml.dom.minidom.parse('Sample_XML.xml')
children = doc.getElementsByTagName('family')
for child in children:
    print(child.getAttribute('name'))
    print(child.getElementsByTagName('child')[0].childNodes[0].nodeValue)
    print(child.getElementsByTagName('child')[1].childNodes[0].nodeValue)
    print(child.getElementsByTagName('child')[2].childNodes[0].nodeValue)
    print(child.getElementsByTagName('child')[2].childNodes[0].nodeValue)

问题分析

原有代码存在两个核心问题:

  • 硬编码索引导致不灵活:通过固定索引[0][1][2]获取child节点,若XML中child数量变化会直接报错。
  • 无法处理嵌套节点:当child标签内包含grandchildren子元素时,childNodes[0]仅能获取到child的文本内容(如Third),无法定位到grandchildren节点。

解决方案

推荐使用Python内置的xml.etree.ElementTree模块(比xml.dom.minidom更简洁易用),结合pandas实现数据提取与电子表格导出。

完整代码示例

import xml.etree.ElementTree as ET
import pandas as pd

# 解析XML文件
tree = ET.parse('Sample_XML.xml')
root = tree.getroot()

# 初始化数据存储列表
data_list = []

# 遍历每个family节点
for family in root.findall('family'):
    family_name = family.attrib.get('name')
    # 遍历当前family下的所有child节点
    for child in family.findall('child'):
        child_name = child.attrib.get('name')
        # 提取child的文本内容,去除多余空白字符
        child_text = ''.join(child.itertext()).strip()
        # 初始化孙辈数据字段
        data1 = None
        data2 = None
        unique = None
        # 查找当前child下的grandchildren节点
        grandchildren = child.find('grandchildren')
        if grandchildren is not None:
            # 提取data节点内容(取前两个)
            data_nodes = grandchildren.findall('data')
            if len(data_nodes) >=1:
                data1 = data_nodes[0].text.strip()
            if len(data_nodes) >=2:
                data2 = data_nodes[1].text.strip()
            # 提取unique节点内容
            unique_node = grandchildren.find('unique')
            if unique_node is not None:
                unique = unique_node.text.strip()
        # 将数据添加到列表
        data_list.append({
            'Family Name': family_name,
            'Child Name': child_name,
            'Child Order': child_text.split('\n')[0].strip(),  # 只取First/Second这类顺序文本
            'Data 1': data1,
            'Data 2': data2,
            'Unique': unique
        })

# 转换为DataFrame并导出到Excel
df = pd.DataFrame(data_list)
df.to_excel('family_data.xlsx', index=False)
print("数据已成功导出到family_data.xlsx")

代码说明

  • XML解析:使用ElementTree的findall方法灵活遍历节点,避免硬编码索引。
  • 文本处理:通过itertext()获取节点下所有文本内容,并用strip()去除多余空白。
  • 嵌套节点处理:通过find方法定位grandchildren节点,再提取其中的data和unique内容。
  • 表格导出:使用pandas将收集的数据转换为DataFrame,一键导出为Excel文件,符合电子表格格式要求。

内容的提问来源于stack exchange,提问作者Smi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 23:02:36