You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python multiprocessing Pool实现文件哈希及大文件夹重复文件查找?

解决多进程计算共享文件夹文件哈希及重复文件查找问题

第一段代码的问题与修正

你的第一段代码无输出且无报错,核心问题有两个:

  1. apply_async参数传递错误:第二个参数必须是参数元组,你直接传入文件路径,导致函数接收的参数不符合定义要求。
  2. 主进程未等待子进程执行:Pool的上下文管理器会在退出时关闭进程池,但不会等待异步任务完成,主进程提前结束后子进程可能还未启动。

修正后的代码:

from multiprocessing import Pool
import os
import hashlib

def hashfile(fp):
    # 大文件分块计算哈希,避免内存溢出
    hash_obj = hashlib.sha1()
    with open(fp, 'rb') as file:
        while chunk := file.read(4096):  # 每次读取4KB块
            hash_obj.update(chunk)
    hash_value = hash_obj.hexdigest()
    print(fp, hash_value)
    return (fp, hash_value)

def testProcessPool(share_input):
    with Pool(processes=2) as pool:
        tasks = []
        for root, dirs, files in os.walk(share_input):
            for f in files:
                fp = os.path.join(root, f)
                # 正确传递参数:args为元组,单个参数需加逗号
                task = pool.apply_async(hashfile, args=(fp,))
                tasks.append(task)
        # 等待所有异步任务执行完毕
        for task in tasks:
            task.get()

if __name__ == '__main__':
    testProcessPool('/Users/gm/Desktop/test')

第二段代码的完善与优化

第二段代码使用imap_unordered的思路合理,但需要补充get_files函数实现,同时优化哈希计算逻辑,加入按文件大小预过滤(这是提升重复文件查找性能的核心:不同大小的文件不可能重复,可直接跳过哈希计算)。

完整实现代码:

from multiprocessing import Pool
import os
import hashlib

def get_files(share_input):
    """生成所有文件的路径迭代器"""
    for root, dirs, files in os.walk(share_input):
        for f in files:
            yield os.path.join(root, f)

def hashfile(fp):
    """分块计算文件哈希,返回(文件路径, 哈希值, 文件大小)"""
    file_size = os.path.getsize(fp)
    hash_obj = hashlib.sha1()
    try:
        with open(fp, 'rb') as file:
            while chunk := file.read(4096):
                hash_obj.update(chunk)
        return (fp, hash_obj.hexdigest(), file_size)
    except PermissionError:
        # 处理共享文件夹权限限制
        return (fp, None, file_size)

def testProcessPool(share_input):
    # 第一步:按文件大小分组,过滤不可能重复的文件
    size_groups = {}
    with Pool() as pool:  # 默认使用CPU核心数作为进程数,平衡性能与开销
        for fp, _, file_size in pool.imap_unordered(hashfile, get_files(share_input)):
            if file_size not in size_groups:
                size_groups[file_size] = []
            size_groups[file_size].append(fp)
    
    # 第二步:仅对同大小文件计算哈希,查找重复
    duplicate_dict = {}
    with Pool() as pool:
        for size, files in size_groups.items():
            if len(files) < 2:
                continue  # 单个文件无需计算哈希
            for fp, hash_val, _ in pool.imap_unordered(hashfile, files):
                if hash_val is None:
                    continue
                if hash_val not in duplicate_dict:
                    duplicate_dict[hash_val] = [fp]
                else:
                    duplicate_dict[hash_val].append(fp)
    
    # 输出重复文件结果
    for hash_val, paths in duplicate_dict.items():
        if len(paths) >= 2:
            print(f"哈希值: {hash_val}")
            print("重复文件路径:")
            for path in paths:
                print(f"  - {path}")
            print("---")

if __name__ == '__main__':
    testProcessPool('/Users/gm/Desktop/test')

关键性能优化点

  • 分块计算哈希:避免一次性读取大文件占用大量内存,适配150GB级别的共享文件夹处理。
  • 按文件大小预过滤:将文件按大小分组,仅对同大小文件计算哈希,大幅减少不必要的计算量。
  • 使用imap_unordered:比apply_async更简洁,直接迭代返回结果,无需手动管理任务队列。
  • 合理设置进程数:默认使用CPU核心数,避免过多进程导致CPU调度开销。
  • 异常处理:捕获共享文件夹的权限异常,避免程序中途崩溃。

内容的提问来源于stack exchange,提问作者pacopyc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 12:35:21