基于VGG16迁移学习训练CIFAR10模型精度无提升问题排查
VGG16迁移学习训练CIFAR10精度无提升排查
我尝试用VGG16做迁移学习训练CIFAR10模型,但训练轮次推进后精度始终没提升,精度一直在10%左右(和随机猜测效果相当),怀疑是输入环节出了问题。以下是我的代码实现及训练输出:
代码实现
import tensorflow as tf from tensorflow.keras import layers, utils from tensorflow.keras.applications import VGG16 from tensorflow.keras.applications.vgg16 import preprocess_input from tensorflow.keras.optimizers import SGD from tensorflow.keras.preprocessing.image import ImageDataGenerator from tensorflow.keras.datasets import cifar10 import numpy as np def build_model(): input_tensor = layers.Input(shape=(32, 32, 3)) resize_layer = layers.Lambda(lambda image: tf.image.resize(image, (224, 224)))(input_tensor) base_model = VGG16(weights='imagenet', include_top=False, input_tensor=resize_layer) base_model.trainable = False x = layers.Flatten()(base_model.output) x = layers.Dense(1024, activation='relu')(x) x = layers.Dropout(0.5)(x) x = layers.Dense(256)(x) x = layers.Dropout(0.5)(x) output_tensor = layers.Dense(10, activation='softmax')(x) model = tf.keras.models.Model(inputs=input_tensor, outputs=output_tensor) return model def cutout(img): mask_size = 16 img_height, img_width, _ = img.shape x = np.random.randint(0, img_height) y = np.random.randint(0, img_width) x1 = np.clip(x - mask_size // 2, 0, img_width) x2 = np.clip(x + mask_size // 2, 0, img_width) y1 = np.clip(y - mask_size // 2, 0, img_height) y2 = np.clip(y + mask_size // 2, 0, img_height) img[y1:y2, x1:x2, :] = 0 return img def preprocessing(img): if np.random.rand() <= 0.5: img = cutout(img) return preprocess_input(img) (x_train, y_train), (x_test, y_test) = cifar10.load_data() y_train = utils.to_categorical(y_train, 10) y_test = utils.to_categorical(y_test, 10) x_train, x_val = x_train[:40000], x_train[40000:] y_train, y_val = y_train[:40000], y_train[40000:] datagen = ImageDataGenerator( rotation_range=15, width_shift_range=0.1, height_shift_range=0.1, horizontal_flip=True, preprocessing_function = preprocessing) train_gen = datagen.flow(x_train, y_train, batch_size=200) model = build_model() model.compile(optimizer=SGD(momentum = 0.9), loss = 'categorical_crossentropy', metrics = ['accuracy']) history1 = model.fit(train_gen, validation_data=(x_val, y_val), epochs=5) model.trainable = True history2 = model.fit(train_gen, validation_data=(x_val, y_val), epochs=10)
训练输出
Epoch 1/5 200/200 [==============================] - 67s 284ms/step - loss: 162.0445 - accuracy: 0.0984 - val_loss: 63.9608 - val_accuracy: 0.1016 Epoch 2/5 200/200 [==============================] - 55s 273ms/step - loss: 59.8545 - accuracy: 0.0995 - val_loss: 30.2987 - val_accuracy: 0.0980 Epoch 3/5 200/200 [==============================] - 54s 272ms/step - loss: 59.9274 - accuracy: 0.0998 - val_loss: 79.5693 - val_accuracy: 0.0997 Epoch 4/5 200/200 [==============================] - 54s 272ms/step - loss: 58.7012 - accuracy: 0.0974 - val_loss: 25.4963 - val_accuracy: 0.0952 Epoch 5/5 200/200 [==============================] - 54s 272ms/step - loss: 55.8720 - accuracy: 0.1001 - val_loss: 57.3340 - val_accuracy: 0.0977
问题排查与修复方案
1. 预处理顺序错误(核心问题)
你先执行cutout数据增强,再调用preprocess_input,但cutout是在0-255的原始像素值上操作,而preprocess_input会将像素值转换为符合VGG16要求的分布(如减去ImageNet均值,范围变为负数),这会导致cutout的填充值0破坏数据分布,甚至让模型无法学习有效特征。
修复方案:
要么先做预处理再执行cutout:
def preprocessing(img): img = preprocess_input(img) if np.random.rand() <= 0.5: img = cutout(img) return img
要么调整cutout的填充值为preprocess_input对应的均值:
def cutout(img): mask_size = 16 img_height, img_width, _ = img.shape x = np.random.randint(0, img_height) y = np.random.randint(0, img_width) x1 = np.clip(x - mask_size // 2, 0, img_width) x2 = np.clip(x + mask_size // 2, 0, img_width) y1 = np.clip(y - mask_size // 2, 0, img_height) y2 = np.clip(y + mask_size // 2, 0, img_height) # 填充ImageNet RGB均值,匹配VGG预处理逻辑 img[y1:y2, x1:x2, :] = [103.939, 116.779, 123.68] return img
2. Cutout参数不适配小尺寸图像
CIFAR10图像仅32x32,你设置的mask_size=16会遮挡1/4的图像区域,对于小图像来说遮挡面积过大,直接破坏关键特征。建议将mask_size调整为8或更小。
3. 验证集未做一致预处理
训练集通过ImageDataGenerator做了预处理,但验证集x_val直接传入训练,未执行相同的预处理步骤,导致训练/验证数据分布不一致,模型无法正确评估性能。
修复方案:
# 对验证集执行相同预处理 x_val_processed = np.array([preprocess_input(img) for img in x_val]) # 训练时替换验证集 history1 = model.fit(train_gen, validation_data=(x_val_processed, y_val), epochs=5)
4. 优化器学习率过高
冻结VGG16骨干网络时,使用SGD默认学习率0.01过高,会导致损失震荡无法收敛。建议降低学习率:
model.compile(optimizer=SGD(momentum=0.9, lr=0.001), loss='categorical_crossentropy', metrics=['accuracy']) # 或改用Adam优化器 model.compile(optimizer=tf.keras.optimizers.Adam(lr=0.0001), loss='categorical_crossentropy', metrics=['accuracy'])
5. 全连接层参数冗余
VGG16输出7x7x512的特征图,Flatten后是25088维,直接连接1024维Dense层参数跳跃过大,易出现梯度消失。建议改用全局平均池化减少参数:
x = layers.GlobalAveragePooling2D()(base_model.output) x = layers.Dense(256, activation='relu')(x) x = layers.Dropout(0.5)(x) output_tensor = layers.Dense(10, activation='softmax')(x)
内容的提问来源于stack exchange,提问作者kogle
相关产品推荐
相关产品推荐

