You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中读取含重复列标题的CSV并保留重复列数据?

处理CSV重复列标题的读取方案

当使用csv.DictReader读取包含重复列标题的CSV时,重复列的值会被覆盖。要实现重复列值以列表存储、非重复列保持原格式的需求,可以通过csv.reader手动处理表头和行数据:

实现代码

import csv

def read_csv_with_duplicate_cols(file_path):
    with open(file_path, 'r', newline='') as f:
        reader = csv.reader(f)
        # 读取表头并统计每个标题的出现次数
        headers = next(reader)
        header_counts = {h: headers.count(h) for h in headers}
        
        result = []
        for row in reader:
            row_dict = {}
            for idx, header in enumerate(headers):
                value = row[idx]
                if header_counts[header] > 1:
                    # 重复列用列表存储值
                    if header not in row_dict:
                        row_dict[header] = []
                    # 若需保留空值,直接移除下面的if判断
                    if value:
                        row_dict[header].append(value)
                else:
                    # 非重复列直接赋值
                    row_dict[header] = value
            result.append(row_dict)
        return result

示例用法

假设CSV文件内容如下:

id,labels,labels,name
1,tag1,tag2,Alice
2,tag3,,Bob

调用函数后得到的结果:

[
    {'id': '1', 'labels': ['tag1', 'tag2'], 'name': 'Alice'},
    {'id': '2', 'labels': ['tag3'], 'name': 'Bob'}
]

关键思路

  1. 先用csv.reader读取表头,统计每个标题的出现次数,判断哪些是重复列
  2. 逐行处理数据时,对重复列的对应值进行收集,存入列表;非重复列直接赋值
  3. 可根据需求选择是否保留空值(移除代码中的if value:判断即可保留)

内容的提问来源于stack exchange,提问作者dibade89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 23:47:11