神经网络训练集准确率超90%但验证集仅约60%,求优化方案
训练集准确率超90%但验证集仅60%的过拟合问题排查与解决
训练神经网络时出现过拟合:训练集准确率超过90%,但验证集准确率停滞在约60%。已尝试多种归一化方式(均值标准差、0-1、-1-1归一化)、调整隐藏层/神经元数量、更换激活函数(ReLU、Tanh),均未解决问题。
代码重现
import numpy as np import pandas as pd import tensorflow as tf from tensorflow.keras.layers import Dense, Input, Softmax , LeakyReLU from tensorflow.keras.layers import BatchNormalization from tensorflow.keras.models import Model from tensorflow.keras import optimizers import sklearn as sk from sklearn import preprocessing from tensorflow.keras.datasets import mnist import matplotlib.pyplot as plt from scipy import stats n_train = 70 import scipy.io data = scipy.io.loadmat( 'process_dataset.mat' ) data = data['data'] y_train0 = data[:,1] #y labels x_train0 = data[:,4:51] #x labels #encode y_train = y_train0[0:n_train].reshape(-1,) y_train = pd.get_dummies(y_train) #Encode ylabel y_train = np.array( y_train.astype(int) ) y_trainOG = y_train #Get original data to be used later for shuffling #scale data x_train = x_train0[0:n_train] xmean = np.mean(x_train, 0) std = np.std(x_train,0) #scale data xmin = np.min(x_train, 0) xmax = np.max(x_train, 0) a1 = (x_train - xmin) x_train = 2*a1/(xmax - xmin) -1 #encode y_test = y_train0[n_train:n_train+30].reshape(-1,) y_test = pd.get_dummies(y_test) #Encode ylabel y_test = np.array( y_test.astype(int) ) #scale test data x_test = x_train0[n_train:n_train+30] xmin = np.min(x_test, 0) xmax = np.max(x_test, 0) xmean = np.mean( x_test, 0) std = np.std(x_test, 0) x_test = 2*(x_test - xmin)/(xmax - xmin) -1 def my_model(x): x = Dense( 50 )(x) x = LeakyReLU()(x) x = Dense(25 )(x) x = LeakyReLU()(x) x = Dense( 50 )(x) x = LeakyReLU()(x) #x = Dense( 10 )(x) x = Dense(5 , activation='softmax', name='output1')(x) return x main_input = Input((47,)) out1 = my_model( main_input ) model = Model( inputs = [main_input ] , outputs=[out1 ] ) model.summary() optim = optimizers.SGD(lr= 0.001, momentum=0 ) loss1 = tf.keras.losses.categorical_crossentropy losses = { "output1":loss1 } model.compile(optimizer=optim, loss=losses, metrics=['accuracy'] ) history = model.fit( x_train, [y_train ], epochs=2000, batch_size=5 , shuffle = True, validation_data = (x_test, y_test) ) a1 = model.evaluate(x_train, [y_train ]) acc = history.history['accuracy'] plt.plot(acc)
问题解决思路与方案
1. 修复测试集归一化逻辑
测试集不能基于自身的min/max或均值标准差做归一化,必须复用训练集计算得到的归一化参数,否则会破坏数据分布一致性,导致模型泛化失效。
# 训练集归一化(保留计算出的xmin/xmax) x_train = x_train0[0:n_train] xmin = np.min(x_train, 0) xmax = np.max(x_train, 0) x_train = 2*(x_train - xmin)/(xmax - xmin) -1 # 测试集复用训练集参数做归一化 x_test = x_train0[n_train:n_train+30] x_test = 2*(x_test - xmin)/(xmax - xmin) -1
2. 应对小数据集过拟合
训练集仅70样本,测试集30样本,数据量过小是核心问题之一:
- 采用交叉验证替代单一划分:比如5折交叉验证,更准确评估模型真实性能
- 尝试结构化数据增强:给特征添加微小噪声、构造特征组合(如多项式特征)、样本权重调整等
- 简化模型结构:当前3层隐藏层(50+25+50神经元)对小数据集过于复杂,可尝试减少为1-2层隐藏层,或降低神经元数量(如20+10)
3. 添加正则化约束
当前模型无任何正则化,直接针对过拟合添加:
- Dropout层:在隐藏层激活后添加,随机丢弃部分神经元减少依赖
- L2正则化:给Dense层添加权重约束,避免权重过大
- 早停机制:监控验证集损失,停止无效迭代,保留最优模型
from tensorflow.keras.callbacks import EarlyStopping from tensorflow.keras.regularizers import l2 def my_model(x): x = Dense(50, kernel_regularizer=l2(0.001))(x) x = LeakyReLU()(x) x = tf.keras.layers.Dropout(0.2)(x) x = Dense(25, kernel_regularizer=l2(0.001))(x) x = LeakyReLU()(x) x = tf.keras.layers.Dropout(0.2)(x) x = Dense(5, activation='softmax', name='output1')(x) return x # 早停回调 early_stop = EarlyStopping(monitor='val_loss', patience=30, restore_best_weights=True) history = model.fit(..., callbacks=[early_stop])
4. 优化训练配置
当前优化器设置不合理:
- SGD学习率0.001过低且无动量,收敛速度慢且易陷入局部最优,可调整为
SGD(lr=0.01, momentum=0.9) - 更换为Adam优化器:自适应学习率更适合小数据集训练
optim = tf.keras.optimizers.Adam(learning_rate=0.001)
5. 启用Batch Normalization
导入了BatchNormalization但未使用,在隐藏层激活前添加可稳定训练分布,提升泛化能力:
def my_model(x): x = Dense(50)(x) x = BatchNormalization()(x) x = LeakyReLU()(x) x = Dense(25)(x) x = BatchNormalization()(x) x = LeakyReLU()(x) x = Dense(5, activation='softmax', name='output1')(x) return x
内容的提问来源于stack exchange,提问作者gingerorange
相关产品推荐
相关产品推荐

