如何用ElementTree与Python将指定XML解析为Pandas DataFrame?
解析带IEC62325命名空间的GL_MarketDocument XML到Pandas DataFrame的简洁方案
你遇到的核心问题是IEC62325标准的XML命名空间未被正确处理——不管是ElementTree还是Pandas的XML解析工具,默认会忽略命名空间,直接按标签名查找会找不到内容,只能手动遍历提取,导致操作繁琐。下面是两种简洁的解决方法:
方法一:优化ElementTree的批量提取
通过定义命名空间映射,结合XPath表达式一次性批量提取目标字段,无需手动逐个生成列表:
import xml.etree.ElementTree as ET import pandas as pd # 加载XML文件 tree = ET.parse('your_market_document.xml') root = tree.getroot() # 替换为你XML中实际的IEC62325命名空间URI ns_map = {'iec': 'urn:iec62325.351:tc57wg16:451-6:marketdocument:1:0'} # 用XPath批量提取所有目标标签内容 psr_types = [node.text for node in root.findall('.//iec:psrType', ns_map)] positions = [node.text for node in root.findall('.//iec:position', ns_map)] quantities = [node.text for node in root.findall('.//iec:quantity', ns_map)] # 直接构建DataFrame df = pd.DataFrame({ 'psrType': psr_types, 'position': positions, 'quantity': quantities })
方法二:用Pandas read_xml直接解析
Pandas的read_xml支持直接处理带命名空间的XML,只需指定命名空间映射和目标节点的XPath路径:
import pandas as pd # 命名空间映射(需与XML中的URI完全一致) ns_map = {'iec': 'urn:iec62325.351:tc57wg16:451-6:marketdocument:1:0'} # 读取XML并解析为DataFrame # 注意:xpath需指向包含psrType、position、quantity的重复父节点(比如每个数据行对应的节点) df = pd.read_xml( 'your_market_document.xml', namespaces=ns_map, xpath='.//iec:Point', # 示例:假设目标字段都在Point节点下 usecols=['psrType', 'position', 'quantity'] )
关键注意事项
- 命名空间URI必须和XML根节点中声明的完全一致,比如不同版本的IEC62325标准可能会有版本号差异,需仔细核对。
- XPath路径要准确对应标签的嵌套层级,比如字段嵌套在
TimeSeries/Period/Point下,就需要写.//iec:TimeSeries/iec:Period/iec:Point/iec:psrType。 - 使用
read_xml时,XPath必须指向重复的数据节点(即每一行数据对应的父节点),否则Pandas无法正确识别行结构。
内容的提问来源于stack exchange,提问作者ruthpozuelo
相关产品推荐
相关产品推荐

