Python 3多CSV字符串导入主CSV并去重及写入问题求助
多CSV文件去重合并到主CSV的解决方案
没问题!我来帮你搞定这个需求~ 下面我会一步步讲清楚怎么实现去重校验和写入主CSV,附上完整可运行的代码,再拆解关键细节。
核心思路
- 先读取主CSV里已有的内容,把每条记录存到一个集合里(集合的查找效率极高,适合快速查重)
- 遍历所有待导入的CSV文件,逐行检查记录是否已存在,不存在就收集起来
- 把新收集的记录追加到主CSV中,完全不会覆盖原有内容
完整代码示例
import csv import os # 配置路径(根据你的实际情况修改) MAIN_CSV_PATH = "main.csv" INPUT_CSV_FOLDER = "your_csv_folder" # 存放所有待导入CSV的文件夹 def load_existing_records(): """加载主CSV中已有的记录,返回表头和用于查重的集合""" existing_records = set() # 如果主CSV已经存在,就读取其中的内容 if os.path.exists(MAIN_CSV_PATH): with open(MAIN_CSV_PATH, 'r', newline='', encoding='utf-8') as f: reader = csv.reader(f) header = next(reader) # 跳过表头 for row in reader: # 把每行转成元组(列表不可哈希,无法存入集合) existing_records.add(tuple(row)) return header, existing_records # 如果主CSV不存在,返回空表头和空集合 return None, existing_records def collect_new_records(existing_records, header): """遍历待导入的CSV,收集所有不存在的新记录""" new_records = [] for filename in os.listdir(INPUT_CSV_FOLDER): if not filename.endswith(".csv"): continue # 跳过非CSV文件 file_path = os.path.join(INPUT_CSV_FOLDER, filename) with open(file_path, 'r', newline='', encoding='utf-8') as f: reader = csv.reader(f) current_header = next(reader) # 第一次运行(主CSV为空)时,记录表头 if header is None: header = current_header # 校验列结构是否一致,避免格式混乱 elif current_header != header: print(f"警告:{filename}的表头与主CSV不一致,已跳过该文件") continue # 逐行检查是否重复 for row in reader: row_tuple = tuple(row) if row_tuple not in existing_records: new_records.append(row) existing_records.add(row_tuple) # 加入集合,避免后续文件重复导入 return header, new_records def write_to_main_csv(header, new_records): """把新记录写入主CSV""" # 打开主CSV,用追加模式('a'),避免覆盖原有内容 with open(MAIN_CSV_PATH, 'a', newline='', encoding='utf-8') as f: writer = csv.writer(f) # 如果主CSV之前不存在,先写入表头 if not os.path.exists(MAIN_CSV_PATH): writer.writerow(header) # 写入所有新记录 writer.writerows(new_records) print(f"成功导入{len(new_records)}条新记录到主CSV!") if __name__ == "__main__": header, existing_records = load_existing_records() header, new_records = collect_new_records(existing_records, header) write_to_main_csv(header, new_records)
关键细节解释
- 查重效率:用集合
existing_records存储已有记录,集合的成员检查是O(1)时间复杂度,比遍历列表快得多,适合处理大量数据。 - 记录存储:把每行转成元组
tuple(row)存入集合,因为列表是不可哈希的,不能作为集合元素,而元组是可哈希的。 - 表头处理:自动识别表头,并且校验所有待导入CSV的表头是否和主CSV一致,避免格式混乱。
- 编码与换行:打开文件时指定
encoding='utf-8'避免中文乱码,newline=''是csv模块推荐的写法,防止出现多余空行。 - 追加模式:用
'a'模式打开主CSV,只会在文件末尾追加新内容,绝对不会覆盖原有数据。
自定义调整建议
- 如果只需要某一列不重复(比如只校验"ID"列),可以修改
load_existing_records和collect_new_records里的逻辑,只把该列的值存入集合,比如existing_records.add(row[0])(假设ID是第一列)。 - 如果你的CSV用了非逗号分隔符(比如制表符),可以在
csv.reader和csv.writer里指定delimiter='\t'参数。
内容的提问来源于stack exchange,提问作者Veeshi McMillan
相关产品推荐
相关产品推荐

