You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将含多子元素的XML解析为DataFrame:实现每个Fruit对应4行数据

修正XML解析代码实现每个对应多行DataFrame

示例XML结构

假设你的XML文件结构如下(每个包含4个需拆分的子元素):

<Fruits>
    <Fruit>
        <Name>Apple</Name>
        <Color>Red</Color>
        <Price season="Spring">1.2</Price>
        <Price season="Summer">1.0</Price>
        <Price season="Autumn">1.5</Price>
        <Price season="Winter">1.8</Price>
    </Fruit>
    <Fruit>
        <Name>Banana</Name>
        <Color>Yellow</Color>
        <Price season="Spring">0.8</Price>
        <Price season="Summer">0.7</Price>
        <Price season="Autumn">0.9</Price>
        <Price season="Winter">1.1</Price>
    </Fruit>
</Fruits>

原错误代码(仅生成2行)

原代码将每个的所有子元素合并为一行,无法拆分出目标多行:

import xml.etree.ElementTree as ET
import pandas as pd

tree = ET.parse('fruits.xml')
root = tree.getroot()

data = []
for fruit in root.findall('Fruit'):
    name = fruit.find('Name').text
    color = fruit.find('Color').text
    prices = [p.text for p in fruit.findall('Price')]
    seasons = [p.get('season') for p in fruit.findall('Price')]
    data.append({
        'Name': name,
        'Color': color,
        'Spring Price': prices[0],
        'Summer Price': prices[1],
        'Autumn Price': prices[2],
        'Winter Price': prices[3]
    })

df = pd.DataFrame(data)
print(df)

修正后的代码(生成8行)

核心逻辑是嵌套遍历每个需拆分的子元素,将父元素的公共字段与子元素字段组合,逐个添加到数据列表:

import xml.etree.ElementTree as ET
import pandas as pd

tree = ET.parse('fruits.xml')
root = tree.getroot()

data = []
for fruit in root.findall('Fruit'):
    # 提取<Fruit>的公共属性
    name = fruit.find('Name').text
    color = fruit.find('Color').text
    # 遍历每个子元素,生成独立记录
    for price_elem in fruit.findall('Price'):
        season = price_elem.get('season')
        price = price_elem.text
        data.append({
            'Name': name,
            'Color': color,
            'Season': season,
            'Price': price
        })

df = pd.DataFrame(data)
print(df)

关键说明

  • 原问题出在将同一下的所有子元素打包成单条记录,修正后通过嵌套循环,为每个子元素生成独立行。
  • 如果你的XML子元素不是,只需调整fruit.findall('XXX')中的标签名,保持「公共字段+子元素字段」的组合逻辑即可。

内容的提问来源于stack exchange,提问作者Pavel Andreev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 14:15:30