You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用ElementTree与Python将指定XML解析为Pandas DataFrame?

解析带IEC62325命名空间的GL_MarketDocument XML到Pandas DataFrame的简洁方案

你遇到的核心问题是IEC62325标准的XML命名空间未被正确处理——不管是ElementTree还是Pandas的XML解析工具,默认会忽略命名空间,直接按标签名查找会找不到内容,只能手动遍历提取,导致操作繁琐。下面是两种简洁的解决方法:

方法一:优化ElementTree的批量提取

通过定义命名空间映射,结合XPath表达式一次性批量提取目标字段,无需手动逐个生成列表:

import xml.etree.ElementTree as ET
import pandas as pd

# 加载XML文件
tree = ET.parse('your_market_document.xml')
root = tree.getroot()

# 替换为你XML中实际的IEC62325命名空间URI
ns_map = {'iec': 'urn:iec62325.351:tc57wg16:451-6:marketdocument:1:0'}

# 用XPath批量提取所有目标标签内容
psr_types = [node.text for node in root.findall('.//iec:psrType', ns_map)]
positions = [node.text for node in root.findall('.//iec:position', ns_map)]
quantities = [node.text for node in root.findall('.//iec:quantity', ns_map)]

# 直接构建DataFrame
df = pd.DataFrame({
    'psrType': psr_types,
    'position': positions,
    'quantity': quantities
})

方法二:用Pandas read_xml直接解析

Pandas的read_xml支持直接处理带命名空间的XML,只需指定命名空间映射和目标节点的XPath路径:

import pandas as pd

# 命名空间映射(需与XML中的URI完全一致)
ns_map = {'iec': 'urn:iec62325.351:tc57wg16:451-6:marketdocument:1:0'}

# 读取XML并解析为DataFrame
# 注意:xpath需指向包含psrType、position、quantity的重复父节点(比如每个数据行对应的节点)
df = pd.read_xml(
    'your_market_document.xml',
    namespaces=ns_map,
    xpath='.//iec:Point',  # 示例:假设目标字段都在Point节点下
    usecols=['psrType', 'position', 'quantity']
)

关键注意事项

  • 命名空间URI必须和XML根节点中声明的完全一致,比如不同版本的IEC62325标准可能会有版本号差异,需仔细核对。
  • XPath路径要准确对应标签的嵌套层级,比如字段嵌套在TimeSeries/Period/Point下,就需要写.//iec:TimeSeries/iec:Period/iec:Point/iec:psrType。
  • 使用read_xml时,XPath必须指向重复的数据节点(即每一行数据对应的父节点),否则Pandas无法正确识别行结构。

内容的提问来源于stack exchange,提问作者ruthpozuelo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 15:30:44