模型验证准确率96%但测试图像预测效果差,求解决方案
我训练了一个模型,验证准确率达到96%且验证损失极低,但在测试其他图像时,预测准确率远低于验证集准确率。已尝试让验证阶段与测试阶段的图像采用相同参数处理,但问题仍未解决,求解决思路。
附训练代码
directory = '/Users/anastalib/PycharmProjects/pythonProject1/Banana2' img_width, img_height = 100, 100 img_datagen = ImageDataGenerator( validation_split=0.2 , rescale=1. / 255, ) # test_datagen = ImageDataGenerator(rescale=1. / 255) train_generator = img_datagen.flow_from_directory(directory, shuffle=True, batch_size=16, subset='training', target_size=(img_width, img_height)) valid_generator = img_datagen.flow_from_directory(directory, shuffle=False, batch_size=16, subset='validation', target_size=(img_width, img_height)) resnet_model = Sequential() pretrained_model = tf.keras.applications.ResNet50(include_top=False, input_shape=(100, 100, 3), pooling='avg', weights='imagenet') for layer in pretrained_model.layers: layer.trainable = False resnet_model.add(pretrained_model) resnet_model.add(Flatten()) resnet_model.add(Dense(256, activation='relu')) resnet_model.add(Dropout(0.5)) resnet_model.add(Dense(3, activation='softmax')) resnet_model.summary() resnet_model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy']) t1 = time.time() print(datetime.datetime.now()) history = resnet_model.fit_generator(train_generator, validation_data=valid_generator, steps_per_epoch=train_generator.n // train_generator.batch_size, validation_steps=valid_generator.n // valid_generator.batch_size, epochs=5) print("Training took %s seconds" % (time.time() - t1)) path = 'overripe.jpg' img = tf.keras.utils.load_img(path, target_size=(100, 100)) x = tf.keras.utils.img_to_array(img) # Rescale image. x = x / 255. x = np.expand_dims(x, axis=0) images = np.vstack([x]) classes = resnet_model.predict(images, batch_size=10) print(np.argmax(classes))
具体解决思路
排查测试数据与验证集的分布差异
验证集是从训练目录自动拆分的,可能和你手动测试的图像在拍摄环境(光照、背景)、目标特征(香蕉成熟度细分、损伤情况)上差异极大。比如训练/验证集都是同一场景的香蕉,而测试图是不同光照、角度下的样本,模型从未学习过这类特征,自然泛化失效。建议统计测试图像的类别分布、视觉特征,和训练/验证集做对比。验证数据拆分的合理性
检查flow_from_directory的自动拆分逻辑,是否因为目录结构问题(比如同一类别图像集中存放),导致验证集恰好选中了易分类的样本,无法代表真实数据分布。可以改为手动划分独立的验证集和测试集,彻底避免拆分偏差。提升模型训练的鲁棒性
当前仅训练5个epoch,且冻结了ResNet50所有层,仅训练顶层全连接层,特征提取能力受限。可尝试:- 解冻ResNet50的高层(比如最后10-20层)进行微调,让模型学习更贴合香蕉分类的专属特征
- 增加训练epoch,同时监控训练/验证损失曲线,避免过拟合
- 在
ImageDataGenerator中加入数据增强,提升模型对图像变化的适应能力:img_datagen = ImageDataGenerator( validation_split=0.2, rescale=1. / 255, rotation_range=20, width_shift_range=0.2, height_shift_range=0.2, horizontal_flip=True, brightness_range=[0.8, 1.2] )
统一预处理逻辑
ResNet50在ImageNet训练时的标准预处理不是简单除以255,而是用tf.keras.applications.resnet.preprocess_input将像素值缩放到[-1,1]区间。你当前训练用rescale=1./255,测试也用同样逻辑,但和预训练模型的输入分布不匹配,会导致泛化差。建议统一预处理:- 训练时去掉
rescale=1./255,改用preprocessing_function=tf.keras.applications.resnet.preprocess_input - 测试图像也用相同处理:
x = tf.keras.applications.resnet.preprocess_input(x)
- 训练时去掉
批量评估测试集准确率
不要仅测试单张图像,收集一批带标注的测试图像,用类似验证集的方式批量加载评估,统计整体准确率,避免单张样本的偶然性(比如overripe.jpg本身属于边缘情况或标注错误)。缓解过拟合风险
验证准确率高可能是模型记住了训练样本的细节而非通用特征,除了数据增强,还可尝试:- 提高Dropout比例(比如调到0.6)
- 减少全连接层神经元数量(比如从256降到128)
- 给全连接层添加L2正则化:
Dense(256, activation='relu', kernel_regularizer=tf.keras.regularizers.l2(0.01))
内容的提问来源于stack exchange,提问作者Anas

