如何修正代码实现多XML文件标签/子标签值存入单个CSV
问题修正:读取多XML文件的标签数据到CSV
问题根源
你的代码存在两个核心问题:
- 每个
<Event>标签下的子属性循环结束后才写入CSV,导致每个Event仅保留最后一个子属性的数值 - 冗余的
with open(files,"rb")语句,ET.parse()可直接读取文件路径,无需额外打开文件
修正后的代码
import csv import xml.etree.ElementTree as ET mypath = '/content/drive/MyDrive/Colab Notebooks/CTIDataset' # 用with语句管理CSV文件,自动处理资源释放 with open("/content/drive/MyDrive/ColabNotebooks/TESTCTI.csv",'w',encoding='utf-8', newline='') as csvfile: csvfile_writer = csv.writer(csvfile) csvfile_writer.writerow(["Event_ID","Date","Info","Category","Comment","Value","Type","Att_ID"]) for files in filename: tree = ET.parse(files) root = tree.getroot() # 遍历每个文件中的所有Event标签 for event_tag in root.findall("Event"): ev_id = getattr(event_tag.find("id"), 'text', None) date = getattr(event_tag.find("date"), 'text', None) info = getattr(event_tag.find("info"), 'text', None) # 遍历当前Event下的所有子属性标签,每遍历一个就写入一行 for attribute in event_tag: if attribute: category = getattr(attribute.find("category"), 'text', None) comment = getattr(attribute.find("comment"), 'text', None) value = getattr(attribute.find("value"), 'text', None) type_ = getattr(attribute.find("type"), 'text', None) att_id = getattr(attribute.find("id"), 'text', None) # 将writerow移到内部循环,每个子属性对应一行数据 csv_line = [ev_id, date, info, category, comment, value, type_, att_id] csvfile_writer.writerow(csv_line)
关键修改点
- 移动写入位置:将
csvfile_writer.writerow(csv_line)移入for attribute in event_tag循环内,确保每个子属性都生成独立的CSV行 - 优化文件管理:用
with语句处理CSV文件,自动关闭文件,避免资源泄漏 - 移除冗余代码:删除无意义的
with open(files,"rb")代码块 - 变量名规范:调整变量名(如
Ev_ID改为ev_id)符合Python命名习惯
内容的提问来源于stack exchange,提问作者M Saad
相关产品推荐
相关产品推荐

