You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行CNN时报PIL.UnidentifiedImageError,如何跳过异常或定位问题图片

解决方案

1. 找到问题图片的文件路径

运行以下脚本可遍历训练集、测试集所有图片,直接输出所有损坏无法识别的图片路径:

import os
from PIL import Image

def check_broken_images(root_path):
    broken_imgs = []
    # 遍历目录下所有文件
    for subdir, _, files in os.walk(root_path):
        for file in files:
            # 过滤非图片格式文件
            if file.lower().endswith(('.png', '.jpg', '.jpeg', '.bmp', '.webp')):
                img_path = os.path.join(subdir, file)
                try:
                    img = Image.open(img_path)
                    img.verify() # 验证图片完整性
                except (IOError, SyntaxError, Image.UnidentifiedImageError):
                    broken_imgs.append(img_path)
                    print(f"损坏图片:{img_path}")
    return broken_imgs

# 替换为你自己的训练、测试集路径
train_broken = check_broken_images(train_path)
test_broken = check_broken_images(test_path)
print(f"训练集损坏图片数:{len(train_broken)}, 测试集损坏图片数:{len(test_broken)}")

拿到路径后可手动删除或修复对应的损坏图片。

2. 训练时自动跳过异常图片

自定义安全的图片生成器,训练过程中碰到损坏图片会自动跳过,避免中断训练:

from keras_preprocessing.image import ImageDataGenerator
from keras_preprocessing.image.utils import load_img, img_to_array
from tensorflow.keras.utils import to_categorical
import numpy as np

class SafeImageDataGenerator(ImageDataGenerator):
    def _get_batches_of_transformed_samples(self, index_array):
        batch_x, batch_y = [], []
        for idx in index_array:
            try:
                # 正常加载并处理图片
                img = load_img(self.filepaths[idx], target_size=self.target_size, color_mode=self.color_mode)
                x = img_to_array(img, data_format=self.data_format)
                x = self.random_transform(x)
                x = self.standardize(x)
                batch_x.append(x)
                batch_y.append(self.classes[idx])
            except Exception as e:
                print(f"跳过异常图片:{self.filepaths[idx]}, 错误信息:{str(e)}")
                continue
        # 不足batch大小的部分用现有样本补齐
        while len(batch_x) < self.batch_size:
            batch_x.append(batch_x[0])
            batch_y.append(batch_y[0])
        # 格式化输出
        batch_x = np.array(batch_x, dtype=self.dtype)
        batch_y = np.array(batch_y)
        if self.class_mode == "categorical":
            batch_y = to_categorical(batch_y, num_classes=self.num_classes)
        return batch_x, batch_y

使用时把原来的ImageDataGenerator替换为SafeImageDataGenerator即可,其余参数不用修改。

3. 兼容异常图片的处理代码

如果不想跳过或删除异常图片,可使用多库兼容的加载逻辑,大部分格式异常、后缀名不对的图片都能被正常读取:

import cv2
from PIL import Image

def compatible_load_img(img_path, target_size):
    # 优先用PIL加载
    try:
        img = Image.open(img_path).convert("RGB")
        return img.resize(target_size)
    except:
        # PIL加载失败尝试用OpenCV加载
        try:
            img_cv = cv2.imread(img_path)
            img_cv = cv2.cvtColor(img_cv, cv2.COLOR_BGR2RGB)
            return Image.fromarray(img_cv).resize(target_size)
        # 都加载失败返回全黑占位图,可根据需求调整为其他默认值
        except:
            return Image.new("RGB", target_size, (0,0,0))

将该函数替换到自定义生成器的图片加载逻辑中,即可兼容绝大多数异常图片场景。

内容的提问来源于stack exchange,提问作者liatkatz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 12:45:01