如何批量对比不同文件夹中同名文本文件首列并迁移差异文件
批量对比文件夹文件并迁移差异文件的实现方法
需求说明
现有两个文件夹folder1和folder2,各包含1000个同名文本文件,每个文件内有多行数字数据。需要对比同文件名文件的第一列数据:若两个文件的第一列存在差异(元素不同,无论顺序),则将folder1中的对应文件迁移至output文件夹。
批量实现思路
原单文件代码用difflib做全量文本对比效率较低,我们只需聚焦每行第一列的差异,结合文件系统操作实现批量处理:
- 遍历
folder1中的所有文件 - 检查
folder2中是否存在同名文件 - 提取两个文件每行的第一列数据并对比差异
- 若存在差异,将
folder1中的文件移动到output文件夹
完整代码实现
import os import shutil # 替换为实际文件夹路径 folder1_path = "./folder1" folder2_path = "./folder2" output_path = "./output" # 创建output文件夹(不存在则自动创建) os.makedirs(output_path, exist_ok=True) # 遍历folder1中的所有文件 for filename in os.listdir(folder1_path): # 仅处理文本文件(可根据需求调整后缀) if not filename.endswith(".txt"): continue file1_full_path = os.path.join(folder1_path, filename) file2_full_path = os.path.join(folder2_path, filename) # 跳过folder2中不存在的文件(若有需要) if not os.path.exists(file2_full_path): continue # 提取文件的第一列数据 def get_first_column(file_path): first_cols = set() with open(file_path, "r", encoding="utf-8") as f: for line in f: line = line.strip() # 跳过空行 if not line: continue # 提取第一列(处理可能的格式异常) try: first_col = line.split()[0] first_cols.add(first_col) except IndexError: # 若行无有效内容,跳过 continue return first_cols cols1 = get_first_column(file1_full_path) cols2 = get_first_column(file2_full_path) # 对比第一列是否存在差异 if cols1 != cols2: # 迁移文件到output文件夹 shutil.move(file1_full_path, os.path.join(output_path, filename)) print(f"已迁移差异文件:{filename}")
关键说明
- 路径处理:用
os.path.join拼接路径,避免跨系统路径格式问题 - 异常防护:添加空行、无效行过滤逻辑,防止因文件格式问题报错
- 差异判断:用集合存储第一列数据,快速对比元素是否完全一致;若要求行顺序也一致,可改为逐行对比第一列
- 文件夹创建:
os.makedirs的exist_ok=True参数确保重复运行时不会报错
内容的提问来源于stack exchange,提问作者Happypumpkin pm
相关产品推荐
相关产品推荐

