You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用原生Python对比CSV首列去重并修复代码问题

修复CSV导入函数的两个问题

问题根源分析

  1. 主CSV数据重置丢失:大概率是写入主文件时使用了w(覆盖模式)而非a(追加模式);或是未读取主文件已有ID就直接覆盖写入。
  2. 唯一交易数统计为0:要么是没有正确提取主CSV中的已有ID,要么是待导入CSV的ID匹配逻辑错误,导致无法识别出唯一ID。

修复后的完整代码

def import_unique_transactions(main_csv_path, import_csv_path):
    # 1. 读取主CSV的已有ID,存入集合(集合查找效率远高于列表)
    existing_ids = set()
    header = ""
    try:
        with open(main_csv_path, 'r', encoding='utf-8') as main_file:
            # 读取并保留表头
            header = main_file.readline().strip()
            for line in main_file:
                # 提取首列ID
                parts = line.strip().split(',')
                if parts:
                    existing_ids.add(parts[0])
    except FileNotFoundError:
        # 主CSV不存在时,从待导入文件获取表头并创建主文件
        with open(import_csv_path, 'r', encoding='utf-8') as import_file:
            header = import_file.readline().strip()
        with open(main_csv_path, 'w', encoding='utf-8') as main_file:
            main_file.write(header + '\n')
        existing_ids = set()

    # 2. 处理待导入CSV,统计并追加唯一记录
    unique_count = 0
    with open(import_csv_path, 'r', encoding='utf-8') as import_file:
        # 跳过待导入文件的表头
        import_file.readline()
        with open(main_csv_path, 'a', encoding='utf-8') as main_file:
            for line in import_file:
                line = line.strip()
                if not line:
                    continue
                parts = line.split(',')
                current_id = parts[0]
                if current_id not in existing_ids:
                    main_file.write(line + '\n')
                    unique_count += 1
                    existing_ids.add(current_id)  # 避免同一待导入文件内的重复ID被多次写入

    print(f"成功追加 {unique_count} 条唯一交易记录")
    return unique_count

# 使用示例
import_unique_transactions('transactions_ledgers.csv', 'to_import.csv')

关键修复点说明

  • 防止主CSV被覆盖:写入主文件时使用a(追加)模式,仅在主文件不存在时用w模式创建并写入表头。
  • 正确统计唯一记录:
    • 用集合存储已有ID,O(1)时间复杂度的查找确保效率。
    • 遍历待导入文件时,检查每条记录的ID是否在已有集合中,不在则计数并追加。
    • 追加后将新ID加入集合,避免同一待导入文件内的重复ID被多次写入。
  • 表头处理:主文件不存在时从待导入文件获取表头,避免重复写入表头导致数据格式混乱。

内容的提问来源于stack exchange,提问作者Jorge Daniel Atuesta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 21:25:24