VGG16迁移学习猫狗分类:evaluate与predict结果差异原因排查
猫狗分类模型evaluate与predict准确率差异问题排查
我基于VGG-16通过迁移学习构建了猫狗分类模型,训练完成后使用model.evaluate评估测试集,准确率约为91%;但使用model.predict生成混淆矩阵时,准确率仅约50%。训练及验证过程中准确率均不低于0.8,想了解这种结果差异的原因,或是预测部分代码是否存在错误。
数据生成器代码
trdata = ImageDataGenerator( rescale=1./255, rotation_range=5, zoom_range=(0.95, 0.95), horizontal_flip=True, vertical_flip=True) traindata = trdata.flow_from_directory( directory='/tmp/cats-v-dogs/training/', target_size=(224,224), batch_size=128, shuffle=True, class_mode='categorical') vldata = ImageDataGenerator(rescale=1.0 / 255.) valdata = vldata.flow_from_directory( directory='/tmp/cats-v-dogs/validation/', target_size=(224,224), shuffle=True, class_mode='categorical') testdata = vldata.flow_from_directory( directory='/tmp/cats-v-dogs/testing/', target_size=(224,224), shuffle=True, class_mode='categorical') # 输出: # Found 19998 images belonging to 2 classes. # Found 2500 images belonging to 2 classes. # Found 2500 images belonging to 2 classes.
模型编译代码
from tensorflow.keras import optimizers model_final.compile(loss = "binary_crossentropy", optimizer=optimizers.SGD(learning_rate=0.001, momentum=0.9, decay=0.0005), metrics=["binary_accuracy"]) model_final.summary()
评估结果
test_loss, test_acc= model_final.evaluate(testdata) print(test_loss) print(test_acc) # 输出: # 79/79 [==============================] - 10s 129ms/step - loss: 0.2132 - binary_accuracy: 0.9192 # 0.21318960189819336 # 0.9192000031471252
预测及混淆矩阵结果
# Confution Matrix and Classification Report Y_pred = model_final.predict(testdata, 2500) y_pred = np.argmax(Y_pred, axis=1) print('Confusion Matrix') print(confusion_matrix(testdata.classes, y_pred)) print('Classification Report') target_names = ['Cats', 'Dogs'] print(classification_report(testdata.classes, y_pred, target_names=target_names))
# 输出: Confusion Matrix [[665 585] [685 565]] Classification Report precision recall f1-score support Cats 0.49 0.53 0.51 1250 Dogs 0.49 0.45 0.47 1250 accuracy 0.49 2500 macro avg 0.49 0.49 0.49 2500 weighted avg 0.49 0.49 0.49 2500
问题根源
核心原因是测试集生成器testdata设置了shuffle=True,导致model.predict输出的预测结果顺序与testdata.classes的原始标签顺序完全不匹配:
model.evaluate内部会自动将生成器输出的批次数据和对应标签一一对应计算,即使生成器打乱数据,也能正确匹配每个样本的预测和标签。- 但
testdata.classes存储的是测试集按目录顺序排列的原始标签,而model.predict是按生成器打乱后的顺序输出预测结果,直接将两者传入confusion_matrix会导致标签和预测完全错位,最终得到类似随机猜测的准确率。
修复方案
方案1:关闭测试集生成器的shuffle
重新创建测试集生成器时设置shuffle=False,确保生成器输出顺序与testdata.classes一致:
testdata = vldata.flow_from_directory( directory='/tmp/cats-v-dogs/testing/', target_size=(224,224), shuffle=False, # 关键修改:关闭打乱 class_mode='categorical')
之后重新运行预测和混淆矩阵代码,就能得到正确的结果。
方案2:记录生成器的输出顺序(适用于必须shuffle的场景)
如果需要保持测试集shuffle,可以通过生成器的filenames属性和预测结果对应,手动对齐标签和预测:
# 获取预测结果 Y_pred = model_final.predict(testdata) y_pred = np.argmax(Y_pred, axis=1) # 获取生成器输出的样本对应的真实标签(按生成顺序) y_true = testdata.classes[testdata.index_array] # 生成混淆矩阵 print('Confusion Matrix') print(confusion_matrix(y_true, y_pred)) print('Classification Report') target_names = ['Cats', 'Dogs'] print(classification_report(y_true, y_pred, target_names=target_names))
testdata.index_array存储了生成器当前打乱后的样本索引顺序,用它可以从原始classes中提取对应顺序的真实标签。
额外注意点
模型使用binary_crossentropy损失,但class_mode='categorical'对应输出是2类的softmax输出,这种搭配虽然能运行,但更规范的做法是:
- 如果用
binary_crossentropy,可以将class_mode设为'binary',模型最后一层用Dense(1, activation='sigmoid'),并使用accuracy或binary_accuracy。 - 如果保持
class_mode='categorical',建议改用categorical_crossentropy损失,这样更匹配多分类的输出设置。不过这不是导致本次准确率差异的直接原因,只是优化建议。
内容的提问来源于stack exchange,提问作者Collander
相关产品推荐
相关产品推荐

