运行CNN时报PIL.UnidentifiedImageError,如何跳过异常或定位问题图片
解决方案
1. 找到问题图片的文件路径
运行以下脚本可遍历训练集、测试集所有图片,直接输出所有损坏无法识别的图片路径:
import os from PIL import Image def check_broken_images(root_path): broken_imgs = [] # 遍历目录下所有文件 for subdir, _, files in os.walk(root_path): for file in files: # 过滤非图片格式文件 if file.lower().endswith(('.png', '.jpg', '.jpeg', '.bmp', '.webp')): img_path = os.path.join(subdir, file) try: img = Image.open(img_path) img.verify() # 验证图片完整性 except (IOError, SyntaxError, Image.UnidentifiedImageError): broken_imgs.append(img_path) print(f"损坏图片:{img_path}") return broken_imgs # 替换为你自己的训练、测试集路径 train_broken = check_broken_images(train_path) test_broken = check_broken_images(test_path) print(f"训练集损坏图片数:{len(train_broken)}, 测试集损坏图片数:{len(test_broken)}")
拿到路径后可手动删除或修复对应的损坏图片。
2. 训练时自动跳过异常图片
自定义安全的图片生成器,训练过程中碰到损坏图片会自动跳过,避免中断训练:
from keras_preprocessing.image import ImageDataGenerator from keras_preprocessing.image.utils import load_img, img_to_array from tensorflow.keras.utils import to_categorical import numpy as np class SafeImageDataGenerator(ImageDataGenerator): def _get_batches_of_transformed_samples(self, index_array): batch_x, batch_y = [], [] for idx in index_array: try: # 正常加载并处理图片 img = load_img(self.filepaths[idx], target_size=self.target_size, color_mode=self.color_mode) x = img_to_array(img, data_format=self.data_format) x = self.random_transform(x) x = self.standardize(x) batch_x.append(x) batch_y.append(self.classes[idx]) except Exception as e: print(f"跳过异常图片:{self.filepaths[idx]}, 错误信息:{str(e)}") continue # 不足batch大小的部分用现有样本补齐 while len(batch_x) < self.batch_size: batch_x.append(batch_x[0]) batch_y.append(batch_y[0]) # 格式化输出 batch_x = np.array(batch_x, dtype=self.dtype) batch_y = np.array(batch_y) if self.class_mode == "categorical": batch_y = to_categorical(batch_y, num_classes=self.num_classes) return batch_x, batch_y
使用时把原来的ImageDataGenerator替换为SafeImageDataGenerator即可,其余参数不用修改。
3. 兼容异常图片的处理代码
如果不想跳过或删除异常图片,可使用多库兼容的加载逻辑,大部分格式异常、后缀名不对的图片都能被正常读取:
import cv2 from PIL import Image def compatible_load_img(img_path, target_size): # 优先用PIL加载 try: img = Image.open(img_path).convert("RGB") return img.resize(target_size) except: # PIL加载失败尝试用OpenCV加载 try: img_cv = cv2.imread(img_path) img_cv = cv2.cvtColor(img_cv, cv2.COLOR_BGR2RGB) return Image.fromarray(img_cv).resize(target_size) # 都加载失败返回全黑占位图,可根据需求调整为其他默认值 except: return Image.new("RGB", target_size, (0,0,0))
将该函数替换到自定义生成器的图片加载逻辑中,即可兼容绝大多数异常图片场景。
内容的提问来源于stack exchange,提问作者liatkatz
相关产品推荐
相关产品推荐

