You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VGG16迁移学习猫狗分类:evaluate与predict结果差异原因排查

猫狗分类模型evaluate与predict准确率差异问题排查

我基于VGG-16通过迁移学习构建了猫狗分类模型,训练完成后使用model.evaluate评估测试集,准确率约为91%;但使用model.predict生成混淆矩阵时,准确率仅约50%。训练及验证过程中准确率均不低于0.8,想了解这种结果差异的原因,或是预测部分代码是否存在错误。


数据生成器代码

trdata = ImageDataGenerator(
    rescale=1./255,
    rotation_range=5,
    zoom_range=(0.95, 0.95),
    horizontal_flip=True,
    vertical_flip=True)

traindata = trdata.flow_from_directory(
    directory='/tmp/cats-v-dogs/training/', 
    target_size=(224,224),
    batch_size=128,
    shuffle=True,
    class_mode='categorical')

vldata = ImageDataGenerator(rescale=1.0 / 255.)
valdata = vldata.flow_from_directory(
    directory='/tmp/cats-v-dogs/validation/', 
    target_size=(224,224),
    shuffle=True,
    class_mode='categorical')

testdata = vldata.flow_from_directory(
    directory='/tmp/cats-v-dogs/testing/', 
    target_size=(224,224),
    shuffle=True,
    class_mode='categorical')

# 输出:
# Found 19998 images belonging to 2 classes.
# Found 2500 images belonging to 2 classes.
# Found 2500 images belonging to 2 classes.

模型编译代码

from tensorflow.keras import optimizers
model_final.compile(loss = "binary_crossentropy", optimizer=optimizers.SGD(learning_rate=0.001, momentum=0.9, decay=0.0005), metrics=["binary_accuracy"])
model_final.summary()

评估结果

test_loss, test_acc= model_final.evaluate(testdata)
print(test_loss)
print(test_acc)

# 输出:
# 79/79 [==============================] - 10s 129ms/step - loss: 0.2132 - binary_accuracy: 0.9192
# 0.21318960189819336
# 0.9192000031471252

预测及混淆矩阵结果

# Confution Matrix and Classification Report
Y_pred = model_final.predict(testdata, 2500)
y_pred = np.argmax(Y_pred, axis=1)
print('Confusion Matrix')
print(confusion_matrix(testdata.classes, y_pred))
print('Classification Report')
target_names = ['Cats', 'Dogs']
print(classification_report(testdata.classes, y_pred, target_names=target_names))
# 输出:
Confusion Matrix
[[665 585]
 [685 565]]
Classification Report
              precision    recall  f1-score   support

        Cats       0.49      0.53      0.51      1250
        Dogs       0.49      0.45      0.47      1250

    accuracy                           0.49      2500
   macro avg       0.49      0.49      0.49      2500
weighted avg       0.49      0.49      0.49      2500

问题根源

核心原因是测试集生成器testdata设置了shuffle=True,导致model.predict输出的预测结果顺序与testdata.classes的原始标签顺序完全不匹配:

  • model.evaluate内部会自动将生成器输出的批次数据和对应标签一一对应计算,即使生成器打乱数据,也能正确匹配每个样本的预测和标签。
  • 但testdata.classes存储的是测试集按目录顺序排列的原始标签,而model.predict是按生成器打乱后的顺序输出预测结果,直接将两者传入confusion_matrix会导致标签和预测完全错位,最终得到类似随机猜测的准确率。

修复方案

方案1:关闭测试集生成器的shuffle

重新创建测试集生成器时设置shuffle=False,确保生成器输出顺序与testdata.classes一致:

testdata = vldata.flow_from_directory(
    directory='/tmp/cats-v-dogs/testing/', 
    target_size=(224,224),
    shuffle=False,  # 关键修改:关闭打乱
    class_mode='categorical')

之后重新运行预测和混淆矩阵代码,就能得到正确的结果。

方案2:记录生成器的输出顺序(适用于必须shuffle的场景)

如果需要保持测试集shuffle,可以通过生成器的filenames属性和预测结果对应,手动对齐标签和预测:

# 获取预测结果
Y_pred = model_final.predict(testdata)
y_pred = np.argmax(Y_pred, axis=1)

# 获取生成器输出的样本对应的真实标签(按生成顺序)
y_true = testdata.classes[testdata.index_array]

# 生成混淆矩阵
print('Confusion Matrix')
print(confusion_matrix(y_true, y_pred))
print('Classification Report')
target_names = ['Cats', 'Dogs']
print(classification_report(y_true, y_pred, target_names=target_names))

testdata.index_array存储了生成器当前打乱后的样本索引顺序,用它可以从原始classes中提取对应顺序的真实标签。

额外注意点

模型使用binary_crossentropy损失,但class_mode='categorical'对应输出是2类的softmax输出,这种搭配虽然能运行,但更规范的做法是:

  • 如果用binary_crossentropy,可以将class_mode设为'binary',模型最后一层用Dense(1, activation='sigmoid'),并使用accuracy或binary_accuracy。
  • 如果保持class_mode='categorical',建议改用categorical_crossentropy损失,这样更匹配多分类的输出设置。不过这不是导致本次准确率差异的直接原因,只是优化建议。

内容的提问来源于stack exchange,提问作者Collander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 00:15:32