You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按内容对同目录文件分组?filecmp.cmp()仅支持两两比较

高效分组内容相同但文件名不同的文件方案

这问题我之前处理大量文件时也碰到过,直接用filecmp.cmp()两两对比确实效率太低,尤其是1800个文件的场景,完全没必要。下面给你几个实用的方案,从简单高效到严谨验证的都有,适配你的需求:

方案1:基于文件哈希值分组(推荐)

内容完全相同的文件,它们的哈希值(比如MD5、SHA256)肯定一致,而且计算哈希的效率远高于逐行对比。这个方法特别适合你的场景——毕竟只有20个唯一内容,计算哈希后分组几乎瞬间完成。

代码实现

import os
import hashlib

def calculate_file_hash(file_path, chunk_size=4096):
    """分块计算文件哈希,避免大文件占用过多内存"""
    hash_obj = hashlib.sha256()  # SHA256比MD5碰撞概率更低,可选更换
    with open(file_path, 'rb') as f:
        while chunk := f.read(chunk_size):
            hash_obj.update(chunk)
    return hash_obj.hexdigest()

def group_files_by_content(target_dir):
    # 用哈希值作为键,存储对应文件名列表
    hash_to_files = {}
    for filename in os.listdir(target_dir):
        full_path = os.path.join(target_dir, filename)
        # 只处理txt文件,跳过目录
        if os.path.isfile(full_path) and filename.endswith('.txt'):
            file_hash = calculate_file_hash(full_path)
            if file_hash not in hash_to_files:
                hash_to_files[file_hash] = []
            hash_to_files[file_hash].append(filename)
    
    # 转换为更直观的结构:文件名组 + 内容预览
    result_groups = []
    for file_list in hash_to_files.values():
        # 取组内第一个文件的内容作为代表
        with open(os.path.join(target_dir, file_list[0]), 'r', encoding='utf-8', errors='ignore') as f:
            content = f.read()
        result_groups.append({
            'files': file_list,
            'content': content
        })
    return result_groups

# 调用示例
txt_dir = "/你的txt文件目录路径"
content_groups = group_files_by_content(txt_dir)

# 打印分组结果
for group in content_groups:
    file_str = f"file({', '.join(group['files'])})"
    print(f"{file_str}: {group['content']}")

为什么推荐这个方法?

  • 效率极高:时间复杂度是O(total_file_size),1800个文件完全不在话下;
  • 避免重复对比:不需要两两比较所有文件,大幅减少计算量;
  • 灵活输出:可以轻松转换成字典、列表或pandas数据帧。

方案2:哈希分组+二次验证(严谨版)

如果担心极端情况下的哈希碰撞(概率极低,但确实存在),可以在哈希分组后,对每个组内的文件用filecmp.cmp()做二次验证,确保内容完全一致:

import filecmp

def validate_group_content(group, target_dir):
    """验证组内所有文件内容是否一致"""
    first_file = os.path.join(target_dir, group['files'][0])
    for filename in group['files'][1:]:
        current_file = os.path.join(target_dir, filename)
        if not filecmp.cmp(first_file, current_file, shallow=False):
            # 若发现不一致,拆分到新组(实际场景中几乎不会触发)
            return False, [{'files': [filename], 'content': open(current_file).read()}]
    return True, []

# 对之前的分组结果做验证
final_groups = []
for group in content_groups:
    is_valid, split_groups = validate_group_content(group, txt_dir)
    if is_valid:
        final_groups.append(group)
    else:
        final_groups.extend(split_groups)

转换为DataFrame(可选)

如果你需要用表格形式展示结果,可以把分组转成pandas数据帧:

import pandas as pd

df = pd.DataFrame(content_groups)
# 将文件名列表转为逗号分隔的字符串,方便查看
df['files'] = df['files'].apply(lambda x: ', '.join(x))
print(df)

注意事项

  • 编码问题:如果你的txt文件有不同编码(比如GBK、UTF-8),读取时可以指定encoding参数,或者用errors='ignore'忽略编码错误;
  • 大文件处理:如果文件很大,内容预览可以只取前几行(比如f.read(100)),避免输出过多内容;
  • 性能优化:如果文件数量特别多,可以多线程计算哈希,但你的场景1800个文件单线程足够快了。

内容的提问来源于stack exchange,提问作者sam o

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:46:47