如何使用Python对比两个CSV文件指定列的MD5值并查找重复数据
解决方法
问题分析
你之前的代码问题在于用zip()逐行配对两个文件的行,只能对比相同行号的指定列,完全不符合「跨所有行找两列共同值」的需求。我们可以用集合存储其中一列的所有值,再遍历另一列匹配,效率高逻辑也简单。
实现代码
import csv # 配置部分,按实际情况修改 file1_path = "/root/file1.csv" file2_path = "/root/file2.csv" skip_header = True # 如果你的csv第一行是表头就设为True,没有表头就设为False # 第一步:读取file1的第三列(索引为2)所有非空值存到集合里 file1_md5_set = set() with open(file1_path, 'r', encoding='utf-8') as f: reader = csv.reader(f) if skip_header: next(reader) # 跳过表头行 for row in reader: # 确保行长度足够,且第三列不是空值 if len(row) >=3 and row[2].strip(): file1_md5_set.add(row[2].strip()) # 第二步:读取file2的第一列(索引为0),找共同值 duplicate_md5 = [] with open(file2_path, 'r', encoding='utf-8') as f: reader = csv.reader(f) if skip_header: next(reader) # 跳过表头行 for row in reader: if len(row) >=1 and row[0].strip(): current_md5 = row[0].strip() if current_md5 in file1_md5_set: duplicate_md5.append(current_md5) # 输出结果 print("两列共同存在的md5值如下:") for md5 in duplicate_md5: print(md5) # 如果需要把结果存到文件可以取消注释下面的代码 # with open("duplicate_result.txt", 'w', encoding='utf-8') as f: # f.write('\n'.join(duplicate_md5))
代码说明
- 用集合存file1的md5值是因为集合的成员查询速度远快于列表,就算两个文件有几十万行也不会卡顿
- 加了
strip()处理值前后的空格、换行符,避免因为不可见字符导致匹配失败 - 做了行长度判断,避免短行报索引越界错误
- 支持跳过表头,按你自己的csv实际情况修改
skip_header参数即可
内容的提问来源于stack exchange,提问作者Alex Rebell
相关产品推荐
相关产品推荐

