You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

加载预训练权重后Keras模型(InceptionV3/ResNet121)评估结果异常

问题:加载训练好的InceptionV3/ResNet121权重后评估结果异常,接近初始权重表现

成功训练模型并通过Checkpoint保存权重后,使用load_weights加载权重执行评估,结果却如同加载了初始权重。已在训练集和验证集上测试,排除测试集问题,该问题仅出现在InceptionV3和ResNet121网络,VGG16运行正常。

训练核心代码

def create_inception(model_name, fold_path, model_path, optimizer=Adam(learning_rate=0.0001)):
  inputs = tf.keras.Input(shape=(224, 224, 3))
  head_model = InceptionV3(weights = 'imagenet', include_top = False, input_shape = (224,224,3))

  head_model.trainable = True

  # 关键问题点:固定设置training=True
  head_model = head_model(inputs, training = True)
  head_model = tf.keras.layers.Flatten()(head_model)
  head_model = tf.keras.layers.Dense(256, activation='relu')(head_model)

  output = Dense(3, activation='softmax')(head_model)
  model4 = Model(inputs=inputs, outputs = output)

  # 数据生成器与编译部分省略...
  
  return model4, train_generator, validation_generator

训练结果(正常)

Epoch 40: val_f1_score did not improve from 0.94413
99/99 [==============================] - 54s 540ms/step - loss: 0.0601 - f1_score: 0.9750 - accuracy: 0.9750 - val_loss: 0.1700 - val_f1_score: 0.9347 - val_accuracy: 0.9347

评估代码与异常结果

评估代码:

name = 'kfold_model_' + str(0)
model4, _, _ =  create_inception(name, '/content/Fold10', model_path4, opt['type'](learning_rate=lr))
model4.compile(loss="categorical_crossentropy", 
              optimizer=Adam(learning_rate=0.01),
              metrics=[tfa.metrics.F1Score(num_classes=3, average='micro'), 'accuracy'])
model4.load_weights(model_path4 + "{}_best.hdf5".format(name))
# 测试生成器部分省略...

test_lost, test_f1, test_acc = model4.evaluate(test_generator)

异常输出:

Found 530 images belonging to 3 classes.
17/17 [==============================] - 45s 2s/step - loss: 5.2010 - f1_score: 0.5491 - accuracy: 0.5491
Test f1: 0.5490565896034241
Test Accuracy: 0.5490565896034241

问题原因

InceptionV3和ResNet121包含大量BatchNormalization层,这类层的行为依赖于training参数:

  • 当training=True时,层会使用当前批次的均值/方差更新内部统计量;
  • 当training=False时,层会使用训练阶段累积的均值/方差进行推理。

你的代码在构建模型时固定设置了head_model(inputs, training = True),导致模型在推理阶段仍然强制使用训练模式,忽略了训练好的BatchNormalization统计量,同时也会影响Dropout等层的行为,最终导致评估结果异常。而VGG16不含BatchNormalization层,因此不受此问题影响。

解决方案

  • 移除模型构建时的固定training=True参数
    修改create_inception函数中的模型调用部分,让Keras自动根据场景切换training状态:

    # 原代码:
    # head_model = head_model(inputs, training = True)
    # 修改为:
    head_model = head_model(inputs)
    

    Keras会在model.fit()时自动设置training=True,在model.evaluate()/model.predict()时自动设置training=False,确保BatchNormalization等层在推理时使用正确的统计量。

  • 评估时保持编译参数与训练一致(可选但推荐)
    评估代码中优化器的学习率设置为0.01,与训练时的0.0001不符,虽然不影响权重加载,但建议保持一致,避免潜在问题:

    model4.compile(loss="categorical_crossentropy", 
                  optimizer=Adam(learning_rate=0.0001),  # 与训练时一致
                  metrics=[tfa.metrics.F1Score(num_classes=3, average='micro'), 'accuracy'])
    
  • 验证权重是否正确加载(可选)
    可以通过打印某一层的权重均值,对比训练后和加载后的结果,确认权重已正确加载:

    # 训练后打印权重均值
    print("训练后顶层Dense层权重均值:", model4.layers[-1].get_weights()[0].mean())
    # 加载权重后打印
    model4.load_weights(...)
    print("加载权重后顶层Dense层权重均值:", model4.layers[-1].get_weights()[0].mean())
    

    如果两个均值接近,则说明权重已正确加载,问题确实出在模型的推理模式设置上。

内容的提问来源于stack exchange,提问作者Kongol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 20:05:31