You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3多CSV字符串导入主CSV并去重及写入问题求助

多CSV文件去重合并到主CSV的解决方案

没问题!我来帮你搞定这个需求~ 下面我会一步步讲清楚怎么实现去重校验和写入主CSV,附上完整可运行的代码,再拆解关键细节。

核心思路

  1. 先读取主CSV里已有的内容,把每条记录存到一个集合里(集合的查找效率极高,适合快速查重)
  2. 遍历所有待导入的CSV文件,逐行检查记录是否已存在,不存在就收集起来
  3. 把新收集的记录追加到主CSV中,完全不会覆盖原有内容

完整代码示例

import csv
import os

# 配置路径(根据你的实际情况修改)
MAIN_CSV_PATH = "main.csv"
INPUT_CSV_FOLDER = "your_csv_folder"  # 存放所有待导入CSV的文件夹

def load_existing_records():
    """加载主CSV中已有的记录,返回表头和用于查重的集合"""
    existing_records = set()
    # 如果主CSV已经存在,就读取其中的内容
    if os.path.exists(MAIN_CSV_PATH):
        with open(MAIN_CSV_PATH, 'r', newline='', encoding='utf-8') as f:
            reader = csv.reader(f)
            header = next(reader)  # 跳过表头
            for row in reader:
                # 把每行转成元组(列表不可哈希,无法存入集合)
                existing_records.add(tuple(row))
        return header, existing_records
    # 如果主CSV不存在,返回空表头和空集合
    return None, existing_records

def collect_new_records(existing_records, header):
    """遍历待导入的CSV,收集所有不存在的新记录"""
    new_records = []
    for filename in os.listdir(INPUT_CSV_FOLDER):
        if not filename.endswith(".csv"):
            continue  # 跳过非CSV文件
        file_path = os.path.join(INPUT_CSV_FOLDER, filename)
        with open(file_path, 'r', newline='', encoding='utf-8') as f:
            reader = csv.reader(f)
            current_header = next(reader)
            # 第一次运行(主CSV为空)时,记录表头
            if header is None:
                header = current_header
            # 校验列结构是否一致,避免格式混乱
            elif current_header != header:
                print(f"警告:{filename}的表头与主CSV不一致,已跳过该文件")
                continue
            # 逐行检查是否重复
            for row in reader:
                row_tuple = tuple(row)
                if row_tuple not in existing_records:
                    new_records.append(row)
                    existing_records.add(row_tuple)  # 加入集合,避免后续文件重复导入
    return header, new_records

def write_to_main_csv(header, new_records):
    """把新记录写入主CSV"""
    # 打开主CSV,用追加模式('a'),避免覆盖原有内容
    with open(MAIN_CSV_PATH, 'a', newline='', encoding='utf-8') as f:
        writer = csv.writer(f)
        # 如果主CSV之前不存在,先写入表头
        if not os.path.exists(MAIN_CSV_PATH):
            writer.writerow(header)
        # 写入所有新记录
        writer.writerows(new_records)
    print(f"成功导入{len(new_records)}条新记录到主CSV!")

if __name__ == "__main__":
    header, existing_records = load_existing_records()
    header, new_records = collect_new_records(existing_records, header)
    write_to_main_csv(header, new_records)

关键细节解释

  1. 查重效率:用集合existing_records存储已有记录,集合的成员检查是O(1)时间复杂度,比遍历列表快得多,适合处理大量数据。
  2. 记录存储:把每行转成元组tuple(row)存入集合,因为列表是不可哈希的,不能作为集合元素,而元组是可哈希的。
  3. 表头处理:自动识别表头,并且校验所有待导入CSV的表头是否和主CSV一致,避免格式混乱。
  4. 编码与换行:打开文件时指定encoding='utf-8'避免中文乱码,newline=''是csv模块推荐的写法,防止出现多余空行。
  5. 追加模式:用'a'模式打开主CSV,只会在文件末尾追加新内容,绝对不会覆盖原有数据。

自定义调整建议

  • 如果只需要某一列不重复(比如只校验"ID"列),可以修改load_existing_records和collect_new_records里的逻辑,只把该列的值存入集合,比如existing_records.add(row[0])(假设ID是第一列)。
  • 如果你的CSV用了非逗号分隔符(比如制表符),可以在csv.reader和csv.writer里指定delimiter='\t'参数。

内容的提问来源于stack exchange,提问作者Veeshi McMillan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:32:48