You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练基于预训练VGG19的多输入模型是否需用preprocess_input?

问题描述

我用在ImageNet上预训练且移除顶层(从flatten到最后输出层)的VGG19模型,处理来自修改版ModelNet数据集的多幅图像输入——先让VGG19提取每幅图的特征,拼接后送入自定义的20分类顶层网络。模型代码如下:

from keras.applications.vgg19 import VGG19

input_1 = Input(shape=(224, 224, 3), name='image1')              
input_2 = Input(shape=(224, 224, 3), name='image2')              
input_3 = Input(shape=(224, 224, 3), name='image3')              
input_4 = Input(shape=(224, 224, 3), name='image4')              

base_model = VGG19(weights='imagenet', input_shape=(224, 224, 3), include_top=False)
base_model.trainable = False

x1 = base_model(input_1, training=False)
x2 = base_model(input_2, training=False)
x3 = base_model(input_3, training=False)
x4 = base_model(input_4, training=False)

x = concatenate([x1, x2, x3, x4])

x = Flatten()(x)
x = Dense(512, activation='relu')(x)
x = Dropout(0.5)(x) 
x = Dense(256, activation='relu')(x)
outputs = Dense(20, activation='softmax', name='class_out')(x)

model = tf.keras.models.Model([input_1, input_2, input_3, input_4], outputs)
print(model.summary())

目前我用自定义数据加载器把输入图像按x_batch[i,j]/255做归一化后训练,有两个核心疑问:

  1. 训练时必须用from keras.applications.vgg19 import preprocess_input处理输入吗?还是只在预测时用?
  2. 用preprocess_input替换普通归一化后,图像显示异常,这该怎么解决?
解答

一、preprocess_input的使用时机

训练和预测都必须用,原因很直接:

  • VGG19在ImageNet上预训练时,输入就是经过preprocess_input处理的——具体是把RGB通道转成BGR,再分别减去ImageNet数据集的均值[103.939, 116.779, 123.68]。
  • 如果你只用x/255归一化,输入数据的分布和预训练时完全不匹配,会导致VGG19提取的特征偏离它“熟悉”的模式,直接拉低模型性能,甚至让预训练的权重失去作用。
  • 哪怕你冻结了VGG19的权重,输入分布不对,底层特征提取效果也会大打折扣,所以训练和预测的输入处理逻辑必须完全一致。

二、图像显示异常的解决方法

用preprocess_input后图像显示异常是正常的,因为它的处理逻辑不是为了可视化:

  • 处理后像素值会变成[-123.68, 131.32]左右的范围,还调换了RGB和BGR通道顺序,这超出了图像可视化的正常像素范围([0,1]或[0,255]),所以直接显示会出现偏色、失真等问题。
  • 如果要可视化处理后的图像,需要做反向转换,示例代码如下:
import numpy as np

def reverse_preprocess(img):
    # img是经过preprocess_input处理后的数组
    img = img.copy()
    # 把BGR通道转回RGB
    img = img[..., ::-1]
    # 加上预训练时减去的均值
    img[..., 0] += 103.939
    img[..., 1] += 116.779
    img[..., 2] += 123.68
    # 把像素值裁剪到0-255范围,再转成图像常用的uint8类型
    img = np.clip(img, 0, 255).astype(np.uint8)
    return img
  • 注意:反向转换只是为了看图像,训练和预测时还是得用原始的preprocess_input处理输入,不能搞混。

内容的提问来源于stack exchange,提问作者stackersTech101

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 19:54:25