You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用xml.etree.ElementTree解析XML并导入Pandas DataFrame问题求助

XML Bug仓库数据集解析及DataFrame构建解决方案

问题1:提取字段并将同bug的所有file存为单单元格列表

错误原因说明

你之前的三个尝试分别对应三类问题:

  • 直接使用findall('fixedFiles/file')返回的是Element对象列表,不是文件路径文本,所以看起来是空内容
  • 使用findtext('fixedFiles/file')只会匹配第一个符合条件的节点,自然丢失后续file内容
  • 内层遍历file时直接给全局fixedfile追加元素,会导致file列表长度远大于bug数量,无法和summary、description字段对齐

正确实现代码

import pandas as pd 
import xml.etree.ElementTree as ET

# 读取XML文件
xmldoc = ET.parse('dataset.xml')

# 初始化存储列表
summary_list = []
desc_list = []
fixed_files_list = []

# 遍历所有bug节点
for bug in xmldoc.iter(tag='bug'):
    # 提取summary和description
    summary = bug.findtext('buginformation/summary')
    desc = bug.findtext('buginformation/description')
    # 提取当前bug下所有file的文本,组成列表
    files = [file.text for file in bug.iterfind('./fixedFiles/file')]
    # 追加到对应列表
    summary_list.append(summary)
    desc_list.append(desc)
    fixed_files_list.append(files)

# 构建DataFrame
df = pd.DataFrame({
    'summary': summary_list,
    'description': desc_list,
    'fixed_files': fixed_files_list
})

运行后fixed_files列每个单元格都会存储对应bug的所有修复文件路径列表,三个列表长度完全一致,不会报长度不匹配错误。

问题2:筛选包含2或3个file子元素的bug

有两种实现方式,你可以按需选择:

方式1:遍历XML时直接过滤

只在存储时保留符合条件的bug:

import pandas as pd 
import xml.etree.ElementTree as ET

xmldoc = ET.parse('dataset.xml')
summary_list = []
desc_list = []
fixed_files_list = []

for bug in xmldoc.iter(tag='bug'):
    summary = bug.findtext('buginformation/summary')
    desc = bug.findtext('buginformation/description')
    files = [file.text for file in bug.iterfind('./fixedFiles/file')]
    # 新增判断:仅保留file数量为2或3的bug
    if len(files) in (2,3):
        summary_list.append(summary)
        desc_list.append(desc)
        fixed_files_list.append(files)

df_filtered = pd.DataFrame({
    'summary': summary_list,
    'description': desc_list,
    'fixed_files': fixed_files_list
})

方式2:已生成全量DataFrame后过滤

如果已经生成了全量的df,可以新增file数量列后筛选:

# 新增列存储每个bug的修复文件数量
df['file_count'] = df['fixed_files'].apply(len)
# 筛选file数量为2或3的行
df_filtered = df[df['file_count'].isin([2,3])]

内容的提问来源于stack exchange,提问作者Marc Ouédraogo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 10:18:02