如何在Python中读取含重复列标题的CSV并保留重复列数据?
处理CSV重复列标题的读取方案
当使用csv.DictReader读取包含重复列标题的CSV时,重复列的值会被覆盖。要实现重复列值以列表存储、非重复列保持原格式的需求,可以通过csv.reader手动处理表头和行数据:
实现代码
import csv def read_csv_with_duplicate_cols(file_path): with open(file_path, 'r', newline='') as f: reader = csv.reader(f) # 读取表头并统计每个标题的出现次数 headers = next(reader) header_counts = {h: headers.count(h) for h in headers} result = [] for row in reader: row_dict = {} for idx, header in enumerate(headers): value = row[idx] if header_counts[header] > 1: # 重复列用列表存储值 if header not in row_dict: row_dict[header] = [] # 若需保留空值,直接移除下面的if判断 if value: row_dict[header].append(value) else: # 非重复列直接赋值 row_dict[header] = value result.append(row_dict) return result
示例用法
假设CSV文件内容如下:
id,labels,labels,name 1,tag1,tag2,Alice 2,tag3,,Bob
调用函数后得到的结果:
[ {'id': '1', 'labels': ['tag1', 'tag2'], 'name': 'Alice'}, {'id': '2', 'labels': ['tag3'], 'name': 'Bob'} ]
关键思路
- 先用
csv.reader读取表头,统计每个标题的出现次数,判断哪些是重复列 - 逐行处理数据时,对重复列的对应值进行收集,存入列表;非重复列直接赋值
- 可根据需求选择是否保留空值(移除代码中的
if value:判断即可保留)
内容的提问来源于stack exchange,提问作者dibade89
相关产品推荐
相关产品推荐

