如何按内容对同目录文件分组?filecmp.cmp()仅支持两两比较
高效分组内容相同但文件名不同的文件方案
这问题我之前处理大量文件时也碰到过,直接用filecmp.cmp()两两对比确实效率太低,尤其是1800个文件的场景,完全没必要。下面给你几个实用的方案,从简单高效到严谨验证的都有,适配你的需求:
方案1:基于文件哈希值分组(推荐)
内容完全相同的文件,它们的哈希值(比如MD5、SHA256)肯定一致,而且计算哈希的效率远高于逐行对比。这个方法特别适合你的场景——毕竟只有20个唯一内容,计算哈希后分组几乎瞬间完成。
代码实现
import os import hashlib def calculate_file_hash(file_path, chunk_size=4096): """分块计算文件哈希,避免大文件占用过多内存""" hash_obj = hashlib.sha256() # SHA256比MD5碰撞概率更低,可选更换 with open(file_path, 'rb') as f: while chunk := f.read(chunk_size): hash_obj.update(chunk) return hash_obj.hexdigest() def group_files_by_content(target_dir): # 用哈希值作为键,存储对应文件名列表 hash_to_files = {} for filename in os.listdir(target_dir): full_path = os.path.join(target_dir, filename) # 只处理txt文件,跳过目录 if os.path.isfile(full_path) and filename.endswith('.txt'): file_hash = calculate_file_hash(full_path) if file_hash not in hash_to_files: hash_to_files[file_hash] = [] hash_to_files[file_hash].append(filename) # 转换为更直观的结构:文件名组 + 内容预览 result_groups = [] for file_list in hash_to_files.values(): # 取组内第一个文件的内容作为代表 with open(os.path.join(target_dir, file_list[0]), 'r', encoding='utf-8', errors='ignore') as f: content = f.read() result_groups.append({ 'files': file_list, 'content': content }) return result_groups # 调用示例 txt_dir = "/你的txt文件目录路径" content_groups = group_files_by_content(txt_dir) # 打印分组结果 for group in content_groups: file_str = f"file({', '.join(group['files'])})" print(f"{file_str}: {group['content']}")
为什么推荐这个方法?
- 效率极高:时间复杂度是O(total_file_size),1800个文件完全不在话下;
- 避免重复对比:不需要两两比较所有文件,大幅减少计算量;
- 灵活输出:可以轻松转换成字典、列表或pandas数据帧。
方案2:哈希分组+二次验证(严谨版)
如果担心极端情况下的哈希碰撞(概率极低,但确实存在),可以在哈希分组后,对每个组内的文件用filecmp.cmp()做二次验证,确保内容完全一致:
import filecmp def validate_group_content(group, target_dir): """验证组内所有文件内容是否一致""" first_file = os.path.join(target_dir, group['files'][0]) for filename in group['files'][1:]: current_file = os.path.join(target_dir, filename) if not filecmp.cmp(first_file, current_file, shallow=False): # 若发现不一致,拆分到新组(实际场景中几乎不会触发) return False, [{'files': [filename], 'content': open(current_file).read()}] return True, [] # 对之前的分组结果做验证 final_groups = [] for group in content_groups: is_valid, split_groups = validate_group_content(group, txt_dir) if is_valid: final_groups.append(group) else: final_groups.extend(split_groups)
转换为DataFrame(可选)
如果你需要用表格形式展示结果,可以把分组转成pandas数据帧:
import pandas as pd df = pd.DataFrame(content_groups) # 将文件名列表转为逗号分隔的字符串,方便查看 df['files'] = df['files'].apply(lambda x: ', '.join(x)) print(df)
注意事项
- 编码问题:如果你的txt文件有不同编码(比如GBK、UTF-8),读取时可以指定
encoding参数,或者用errors='ignore'忽略编码错误; - 大文件处理:如果文件很大,内容预览可以只取前几行(比如
f.read(100)),避免输出过多内容; - 性能优化:如果文件数量特别多,可以多线程计算哈希,但你的场景1800个文件单线程足够快了。
内容的提问来源于stack exchange,提问作者sam o
相关产品推荐
相关产品推荐

