You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化基于OpenCV与哈希摘要的Python重复图像检测方案

如何优化基于OpenCV与哈希摘要的Python重复图像检测方案

我太懂你现在的糟心处境了——用当前脚本扫硬盘找重复图,不仅慢得让人抓狂,CPU还被拉满到发烫,完全没法干别的。咱们先拆解现有代码里的几个核心瓶颈,再一步步改成高效版本,既能提速又能降低资源占用。


一、替换文件哈希为感知哈希(Perceptual Hash)

你当前用MD5哈希整个文件内容,存在两个致命问题:

  1. 大图片读完整内容算MD5速度极慢;
  2. 如果两张图像素完全相同但元数据(比如Exif信息)不一样,MD5会判定为不同,但其实是重复图。

感知哈希是基于图片视觉内容计算的,更适配图像去重场景,计算速度也快得多。推荐用dHash(差异哈希),实现简单且对微小改动(如压缩、旋转小角度)的识别效果不错:

import imagehash
from PIL import Image

def generate_perceptual_hash(image_path):
    try:
        with Image.open(image_path) as img:
            # 自动缩小为9x8灰度图并计算差异哈希
            return str(imagehash.dhash(img))
    except Exception as e:
        print(f"处理图片 {image_path} 出错: {e}")
        return None

安装依赖:pip install imagehash pillow


二、删掉冗余的逐像素比对

现有代码里,MD5哈希相同后还做了逐像素校验——这完全是多余的!如果是文件内容哈希相同,文件本身就完全一致,像素肯定没有差异;如果换成感知哈希,哈希相同就代表视觉内容一致(哈希碰撞概率低到可以忽略),直接判定为重复即可,这能省掉大量CPU运算。


三、把O(n²)的嵌套比对改成哈希分组

原来的双层循环遍历所有图片对,时间复杂度是O(n²),几千张图就要做几百万次比对,慢得离谱。我们可以用哈希作为键,把相同哈希的图片归为一组,一次遍历就能找出所有重复组,时间复杂度直接降到O(n):

from collections import defaultdict

def find_duplicate_groups(image_list):
    hash_groups = defaultdict(list)
    for img_info in image_list:
        img_hash = img_info['perceptual_hash']
        if img_hash:
            hash_groups[img_hash].append(img_info)
    # 只保留包含重复图片的组
    return [group for group in hash_groups.values() if len(group) > 1]

四、优化文件夹遍历与图片收集逻辑

你当前用递归+np.append收集图片,效率极低:np.append每次都会重新创建数组,耗时巨大;递归遍历也不如os.walk简洁高效。改成下面的写法:

def get_all_image_paths(scan_folder_path):
    image_paths = []
    for root, dirs, files in os.walk(scan_folder_path):
        # 跳过隐藏文件夹和系统临时文件夹
        dirs[:] = [d for d in dirs if not d.startswith(('.', '$'))]
        for file in files:
            # 只处理常见图片格式,避免无效运算
            if file.lower().endswith(('.png', '.jpg', '.jpeg', '.bmp', '.gif')):
                image_paths.append(os.path.join(root, file))
    return image_paths

五、用多进程并行处理,榨干多核CPU

图像哈希计算是纯CPU密集型任务,Python的GIL锁会限制多线程的CPU利用率,用多进程可以绕过GIL,充分利用所有CPU核心,把计算时间压缩到原来的1/N(N为核心数):

from multiprocessing import Pool

def process_single_image(file_path):
    try:
        img_hash = generate_perceptual_hash(file_path)
        if img_hash:
            return {
                'image_name': os.path.basename(file_path),
                'image_path': file_path,
                'perceptual_hash': img_hash
            }
    except (PermissionError, FileNotFoundError) as e:
        print(f"无法访问文件 {file_path}: {e}")
        return None

def get_all_images_parallel(scan_folder_path):
    image_paths = get_all_image_paths(scan_folder_path)
    # 自动根据CPU核心数创建进程池
    with Pool() as pool:
        results = pool.map(process_single_image, image_paths)
    # 过滤掉处理失败的无效</think_never_used_51bce0c785ca2f68081bfa7d91973934></think_never_used_51bce0c785ca2f68081bfa7d91973934></think_never_used_51bce0c785ca2f68081bfa7d91973934></think_never_used_51bce0c785ca2f68081bfa7d91973934>
</think_never_used_51bce0c785ca2f68081bfa7d91973934>

整合后的完整优化代码

把所有优化点整合起来,最终的高效版本如下:

import os
import json
import datetime
from collections import defaultdict
from multiprocessing import Pool
import imagehash
from PIL import Image

def generate_perceptual_hash(image_path):
    try:
        with Image.open(image_path) as img:
            return str(imagehash.dhash(img))
    except Exception as e:
        print(f"处理图片 {image_path} 出错: {e}")
        return None

def process_single_image(file_path):
    try:
        img_hash = generate_perceptual_hash(file_path)
        if img_hash:
            return {
                'image_name': os.path.basename(file_path),
                'image_path': file_path,
                'perceptual_hash': img_hash
            }
    except (PermissionError, FileNotFoundError) as e:
        print(f"无法访问文件 {file_path}: {e}")
        return None

def get_all_image_paths(scan_folder_path):
    image_paths = []
    for root, dirs, files in os.walk(scan_folder_path):
        dirs[:] = [d for d in dirs if not d.startswith(('.', '$'))]
        for file in files:
            if file.lower().endswith(('.png', '.jpg', '.jpeg', '.bmp', '.gif')):
                image_paths.append(os.path.join(root, file))
    return image_paths

def get_all_images_parallel(scan_folder_path):
    image_paths = get_all_image_paths(scan_folder_path)
    with Pool() as pool:
        results = pool.map(process_single_image, image_paths)
    return [res for res in results if res is not None]

def find_duplicate_groups(image_list):
    hash_groups = defaultdict(list)
    for img_info in image_list:
        hash_groups[img_info['perceptual_hash']].append(img_info)
    return [group for group in hash_groups.values() if len(group) > 1]

if __name__ == '__main__':
    start_time = datetime.datetime.now()
    scan_folder_path = "<你的源文件夹路径>"
    duplicate_json_path = "<重复信息保存路径>"
    log_file_path = os.path.join(duplicate_json_path, "all_duplicate_images.json")

    # 创建保存目录
    if not os.path.exists(duplicate_json_path):
        os.makedirs(duplicate_json_path)
    # 删除旧日志(如果存在)
    if os.path.exists(log_file_path):
        os.remove(log_file_path)

    print("开始扫描并处理图片...")
    all_images = get_all_images_parallel(scan_folder_path)
    print(f"共处理 {len(all_images)} 张图片")

    print("查找重复图片组...")
    duplicate_groups = find_duplicate_groups(all_images)

    # 整理成更清晰的输出格式
    output_data = []
    for idx, group in enumerate(duplicate_groups, 1):
        output_data.append({
            "duplicate_group_id": idx,
            "total_duplicates": len(group),
            "images": group
        })

    # 写入JSON文件
    with open(log_file_path, 'w', encoding='utf-8') as f:
        json.dump(output_data, f, indent=4, ensure_ascii=False)

    end_time = datetime.datetime.now()
    print(f"完成!耗时: {end_time - start_time}")
    print(f"重复图片组已保存到 {log_file_path}")

备注:内容来源于stack exchange,提问作者Manish Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 09:19:33