将含多子元素的XML解析为DataFrame:实现每个Fruit对应4行数据
修正XML解析代码实现每个对应多行DataFrame
示例XML结构
假设你的XML文件结构如下(每个
<Fruits> <Fruit> <Name>Apple</Name> <Color>Red</Color> <Price season="Spring">1.2</Price> <Price season="Summer">1.0</Price> <Price season="Autumn">1.5</Price> <Price season="Winter">1.8</Price> </Fruit> <Fruit> <Name>Banana</Name> <Color>Yellow</Color> <Price season="Spring">0.8</Price> <Price season="Summer">0.7</Price> <Price season="Autumn">0.9</Price> <Price season="Winter">1.1</Price> </Fruit> </Fruits>
原错误代码(仅生成2行)
原代码将每个
import xml.etree.ElementTree as ET import pandas as pd tree = ET.parse('fruits.xml') root = tree.getroot() data = [] for fruit in root.findall('Fruit'): name = fruit.find('Name').text color = fruit.find('Color').text prices = [p.text for p in fruit.findall('Price')] seasons = [p.get('season') for p in fruit.findall('Price')] data.append({ 'Name': name, 'Color': color, 'Spring Price': prices[0], 'Summer Price': prices[1], 'Autumn Price': prices[2], 'Winter Price': prices[3] }) df = pd.DataFrame(data) print(df)
修正后的代码(生成8行)
核心逻辑是嵌套遍历每个需拆分的子元素,将父元素的公共字段与子元素字段组合,逐个添加到数据列表:
import xml.etree.ElementTree as ET import pandas as pd tree = ET.parse('fruits.xml') root = tree.getroot() data = [] for fruit in root.findall('Fruit'): # 提取<Fruit>的公共属性 name = fruit.find('Name').text color = fruit.find('Color').text # 遍历每个子元素,生成独立记录 for price_elem in fruit.findall('Price'): season = price_elem.get('season') price = price_elem.text data.append({ 'Name': name, 'Color': color, 'Season': season, 'Price': price }) df = pd.DataFrame(data) print(df)
关键说明
- 原问题出在将同一
下的所有子元素打包成单条记录,修正后通过嵌套循环,为每个子元素生成独立行。 - 如果你的XML子元素不是
,只需调整 fruit.findall('XXX')中的标签名,保持「公共字段+子元素字段」的组合逻辑即可。
内容的提问来源于stack exchange,提问作者Pavel Andreev
相关产品推荐
相关产品推荐

