如何使用Python xml.dom.minidom访问XML子子节点并导出数据到表格
XML数据提取及电子表格导出解决方案
问题背景
给定如下XML结构,需提取数据并导出为指定格式的电子表格,但原有代码无法访问grandchildren标签内容,需解决该问题。
源XML数据
<?xml version="1.0" encoding="UTF-8" standalone="no"?> <root testAttr="testValue"> <family name="Hardwood"> <child name="Jack">First</child> <child name="Rose">Second</child> <child name="Blue Ivy">Third <grandchildren> <data>One</data> <data>Two</data> <unique>Twins</unique> </grandchildren> </child> <child name="Jane">Fourth</child> </family> <family name="Downie"> <child name="Bill">First</child> <child name="Rosie">Second</child> <child name="Edward"> Third </child> <child name="Jane">Fourth</child> </family> </root>
原有尝试代码
import xml.dom.minidom doc=xml.dom.minidom.parse('Sample_XML.xml') children = doc.getElementsByTagName('family') for child in children: print(child.getAttribute('name')) print(child.getElementsByTagName('child')[0].childNodes[0].nodeValue) print(child.getElementsByTagName('child')[1].childNodes[0].nodeValue) print(child.getElementsByTagName('child')[2].childNodes[0].nodeValue) print(child.getElementsByTagName('child')[2].childNodes[0].nodeValue)
问题分析
原有代码存在两个核心问题:
- 硬编码索引导致不灵活:通过固定索引
[0][1][2]获取child节点,若XML中child数量变化会直接报错。 - 无法处理嵌套节点:当
child标签内包含grandchildren子元素时,childNodes[0]仅能获取到child的文本内容(如Third),无法定位到grandchildren节点。
解决方案
推荐使用Python内置的xml.etree.ElementTree模块(比xml.dom.minidom更简洁易用),结合pandas实现数据提取与电子表格导出。
完整代码示例
import xml.etree.ElementTree as ET import pandas as pd # 解析XML文件 tree = ET.parse('Sample_XML.xml') root = tree.getroot() # 初始化数据存储列表 data_list = [] # 遍历每个family节点 for family in root.findall('family'): family_name = family.attrib.get('name') # 遍历当前family下的所有child节点 for child in family.findall('child'): child_name = child.attrib.get('name') # 提取child的文本内容,去除多余空白字符 child_text = ''.join(child.itertext()).strip() # 初始化孙辈数据字段 data1 = None data2 = None unique = None # 查找当前child下的grandchildren节点 grandchildren = child.find('grandchildren') if grandchildren is not None: # 提取data节点内容(取前两个) data_nodes = grandchildren.findall('data') if len(data_nodes) >=1: data1 = data_nodes[0].text.strip() if len(data_nodes) >=2: data2 = data_nodes[1].text.strip() # 提取unique节点内容 unique_node = grandchildren.find('unique') if unique_node is not None: unique = unique_node.text.strip() # 将数据添加到列表 data_list.append({ 'Family Name': family_name, 'Child Name': child_name, 'Child Order': child_text.split('\n')[0].strip(), # 只取First/Second这类顺序文本 'Data 1': data1, 'Data 2': data2, 'Unique': unique }) # 转换为DataFrame并导出到Excel df = pd.DataFrame(data_list) df.to_excel('family_data.xlsx', index=False) print("数据已成功导出到family_data.xlsx")
代码说明
- XML解析:使用
ElementTree的findall方法灵活遍历节点,避免硬编码索引。 - 文本处理:通过
itertext()获取节点下所有文本内容,并用strip()去除多余空白。 - 嵌套节点处理:通过
find方法定位grandchildren节点,再提取其中的data和unique内容。 - 表格导出:使用
pandas将收集的数据转换为DataFrame,一键导出为Excel文件,符合电子表格格式要求。
内容的提问来源于stack exchange,提问作者Smi
相关产品推荐
相关产品推荐

