You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Pandas read_xml处理嵌套XML子元素并转为目标DataFrame

解决方法

首先纠正基础写法错误:pd.DataFrame('/path/to/xml')不是读取XML文件的方法,该构造方法仅支持从内存结构化数据生成DataFrame,无法解析外部文件格式。

针对你这种嵌套层级固定的XML,最稳妥的方案是用Python标准库xml.etree手动解析,逻辑清晰可控,不需要额外依赖,适配性更强:

import xml.etree.ElementTree as ET
import pandas as pd

# 对应XML头部定义的默认命名空间
NS = {"ns": "Trend_x0020_Report"}

# 解析XML文件
tree = ET.parse("替换为你的XML文件路径.xml")
root = tree.getroot()

rows = []
# 遍历所有时间戳分组
for ts_group in root.findall(".//ns:TimestampGroup", NS):
    current_row = {
        "datetime": ts_group.attrib["textbox30"]
    }
    # 遍历当前时间下的所有数据源
    for source_group in ts_group.findall(".//ns:SourcesGroup", NS):
        source_name = source_group.attrib["textbox29"]
        # 提取对应测量值
        value = float(source_group.find(".//ns:Cell", NS).attrib["textbox5"])
        current_row[source_name] = value
    rows.append(current_row)

# 生成目标格式DataFrame
df = pd.DataFrame(rows)

如果需要用pd.read_xml实现,需要指定正确的XPath路径和命名空间,以下是参考写法:

import pandas as pd

NS = {"ns": "Trend_x0020_Report"}
# 先打平读取所有测量节点的关联属性
df = pd.read_xml(
    "替换为你的XML文件路径.xml",
    xpath=".//ns:TimestampGroup/ns:SourcesGroup_Collection/ns:SourcesGroup/ns:MeasurementGroup_Collection/ns:MeasurementGroup/ns:Cell",
    namespaces=NS,
    attrs_only=True,
    elems_only=False
)
# 后续可自行通过pivot透视得到你需要的宽表格式

如果你的XML结构固定,优先推荐第一种方案,不需要额外处理透视逻辑,直接就能得到要求的输出格式。

内容的提问来源于stack exchange,提问作者MattB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 04:24:05