如何快速检测删除200GB图像数据中损坏图片?如何并行化现有Python代码提速?
性能优化核心思路
你的原始代码运行慢主要来自三个可优化的点:
- 单线程串行处理,磁盘IO的等待时间被完全浪费,200GB的小文件场景下串行的IO开销占比极高
- PIL.Image.open默认会加载全量像素数据,实际上校验图片是否损坏仅需验证文件头和结构,不需要解码全图
- pandas的iterrows()遍历效率极低,非必要场景下可以直接处理路径列表,不需要通过DataFrame中转
并行化改造后的完整代码
代码已修正原实现中的变量笔误,同时加入了并行逻辑和快速校验逻辑,性能可以提升10倍以上:
import PIL import os from PIL import ImageFile from tqdm import tqdm from pathlib import Path import concurrent.futures # 禁止加载截断的损坏图片 ImageFile.LOAD_TRUNCATED_IMAGES = False # 并行线程数,根据磁盘性能调整:固态可以设为32-64,机械硬盘建议设为8-16 MAX_WORKERS = 16 def check_and_delete(img_path): """单张图片校验+删除逻辑""" try: with PIL.Image.open(img_path) as img: # 仅校验图片结构,不加载全量像素,速度提升极多 img.verify() return None except Exception: if os.path.exists(img_path): os.remove(img_path) # 返回损坏的路径用于后续统计 return str(img_path) if __name__ == "__main__": root = Path('some_path') data_root = root / 'dataset' # 直接获取所有jpg路径,不需要生成DataFrame,减少开销 # 如果需要保留label数据,可以保留原来的create_df逻辑,最后取df['path'].tolist()即可 sku_dirs = [i for i in data_root.iterdir() if i.is_dir()] all_img_paths = [j for i in sku_dirs for j in i.glob('*.jpg')] print(f"共找到{len(all_img_paths)}张图片,开始校验...") damaged_paths = [] # 线程池并行处理,IO密集场景线程池开销远小于进程池 with concurrent.futures.ThreadPoolExecutor(max_workers=MAX_WORKERS) as executor: # 用tqdm展示并行进度 for result in tqdm(executor.map(check_and_delete, all_img_paths), total=len(all_img_paths)): if result: damaged_paths.append(result) # 批量打印损坏文件,避免并行时打印混乱 print(f"校验完成,共删除{len(damaged_paths)}张损坏图片,损坏文件列表:") for p in damaged_paths: print(p)
额外优化说明
- 如果需要处理png、webp等其他格式,修改
glob('*.jpg')为glob('*.jpg') + glob('*.png')对应格式即可 - 如果你需要保留原有的DataFrame逻辑做后续处理,只需要把
all_img_paths替换为df['path'].tolist()即可 - 机械硬盘场景下不要把MAX_WORKERS开得太大,过高的随机IO并发反而会拖慢整体读取速度
内容的提问来源于stack exchange,提问作者Igor Terekhin
相关产品推荐
相关产品推荐

