You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中从特定结构XML生成目标格式DataFrame

将特定结构XML转换为Pandas DataFrame的高效方法

解决方案思路

核心是通过标签文本定位(而非索引)遍历XML节点,提取每个Information对应的字段,最终整理成DataFrame结构。这种方法更直观且效率更高,尤其适合结构固定的XML。

代码实现

首先确保安装所需库:

pip install lxml pandas

然后编写解析代码:

from lxml import etree
import pandas as pd

# 示例XML字符串(如果是本地文件,用`etree.parse('your_file.xml')`替代解析逻辑)
xml_content = """
<Errors>
    <Amount>3</Amount>
    <Error>
        <Code>405</Code>
        <Information>
            <Count>1</Count>
            <Time>16:18:13</Time>
            <Parameters>
                <Parameter Parameter="Par1" Value="0"/>
                <Parameter Parameter="Par2" Value="1"/>
            </Parameters>
        </Information>
        <Information>
            <Count>2</Count>
            <Time>11:04:54</Time>
            <Parameters>
                <Parameter Parameter="Par1" Value="2"/>
                <Parameter Parameter="Par2" Value="3"/>
            </Parameters>
        </Information>
    </Error>
    <Error>
        <Code>404</Code>
        <Information>
            <Count>1</Count>
            <Time>20:42:48</Time>
            <Parameters>
                <Parameter Parameter="Par1" Value="4"/>
                <Parameter Parameter="Par2" Value="5"/>
            </Parameters>
        </Information>
    </Error>
</Errors>
"""

# 解析XML
root = etree.fromstring(xml_content)
data = []

# 遍历每个Error节点
for error in root.xpath('//Error'):
    error_code = error.findtext('Code')  # 通过标签名直接获取文本
    # 遍历当前Error下的所有Information节点
    for info in error.xpath('./Information'):
        # 提取基础字段
        count = info.findtext('Count')
        time = info.findtext('Time')
        # 提取Parameters中的参数,转为字典
        param_dict = {
            param.get('Parameter'): param.get('Value')
            for param in info.xpath('./Parameters/Parameter')
        }
        # 将数据添加到列表
        data.append({
            'Code': error_code,
            'Count': count,
            'Time': time,
            'Par1': param_dict.get('Par1'),
            'Par2': param_dict.get('Par2')
        })

# 转换为DataFrame
df = pd.DataFrame(data)
print(df)

代码说明

  1. 节点定位:使用xpath和findtext通过标签名称直接查找节点,避免了依赖索引的脆弱操作,代码可读性更强。
  2. 参数处理:将Parameter节点的属性转为字典,直接通过键名(Par1/Par2)取值,无需按索引遍历参数。
  3. 数据整理:每个Information节点对应DataFrame的一行,确保数据结构符合需求。

注意事项

你的示例XML中Parameters的闭合标签写错了(写成了</Parameter>),实际解析时需要修正为</Parameters>,否则会导致XML解析失败。

内容的提问来源于stack exchange,提问作者Fish1996

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 14:35:58