You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

神经网络训练集准确率超90%但验证集仅约60%,求优化方案

训练集准确率超90%但验证集仅60%的过拟合问题排查与解决

训练神经网络时出现过拟合:训练集准确率超过90%,但验证集准确率停滞在约60%。已尝试多种归一化方式(均值标准差、0-1、-1-1归一化)、调整隐藏层/神经元数量、更换激活函数(ReLU、Tanh),均未解决问题。

代码重现

import numpy as np
import pandas as pd
import tensorflow as tf
from tensorflow.keras.layers import Dense, Input, Softmax , LeakyReLU
from tensorflow.keras.layers import BatchNormalization

from tensorflow.keras.models import Model
from tensorflow.keras import optimizers
import sklearn as sk
from sklearn import preprocessing
from tensorflow.keras.datasets import mnist
import matplotlib.pyplot as plt
from scipy import stats


n_train = 70


import scipy.io
data = scipy.io.loadmat( 'process_dataset.mat' )
data = data['data']
y_train0 = data[:,1]  #y labels
x_train0 = data[:,4:51] #x labels


#encode
y_train = y_train0[0:n_train].reshape(-1,)  
y_train = pd.get_dummies(y_train) #Encode ylabel
y_train = np.array( y_train.astype(int) )
y_trainOG = y_train #Get original data to be used later for shuffling

#scale data
x_train = x_train0[0:n_train]
xmean = np.mean(x_train, 0)
std = np.std(x_train,0)


#scale data
xmin = np.min(x_train, 0)
xmax = np.max(x_train, 0)
a1 = (x_train - xmin)
x_train = 2*a1/(xmax - xmin) -1


#encode
y_test = y_train0[n_train:n_train+30].reshape(-1,)  
y_test = pd.get_dummies(y_test) #Encode ylabel
y_test = np.array( y_test.astype(int) )

#scale test data

x_test = x_train0[n_train:n_train+30]
xmin = np.min(x_test, 0)
xmax = np.max(x_test, 0)
xmean = np.mean( x_test, 0)
std = np.std(x_test, 0) 
x_test = 2*(x_test - xmin)/(xmax - xmin) -1



def my_model(x):

    
    x = Dense( 50 )(x)
    x = LeakyReLU()(x)  
   
    x = Dense(25   )(x)
    x = LeakyReLU()(x)
    
    x = Dense( 50  )(x)
    x = LeakyReLU()(x)
    
    #x = Dense( 10  )(x)


    x = Dense(5 , activation='softmax', name='output1')(x)
    
 
    return x


main_input = Input((47,))

out1 = my_model( main_input    )
model = Model( inputs = [main_input    ] , outputs=[out1 ] )  
model.summary()
optim = optimizers.SGD(lr= 0.001, momentum=0 )
loss1 = tf.keras.losses.categorical_crossentropy

losses = { "output1":loss1 }
model.compile(optimizer=optim, loss=losses, metrics=['accuracy']  )
history =  model.fit( x_train, [y_train ], epochs=2000, batch_size=5 , shuffle = True, validation_data = (x_test, y_test) )

a1 = model.evaluate(x_train, [y_train ])
acc = history.history['accuracy']

plt.plot(acc)

问题解决思路与方案

1. 修复测试集归一化逻辑

测试集不能基于自身的min/max或均值标准差做归一化,必须复用训练集计算得到的归一化参数,否则会破坏数据分布一致性,导致模型泛化失效。

# 训练集归一化(保留计算出的xmin/xmax)
x_train = x_train0[0:n_train]
xmin = np.min(x_train, 0)
xmax = np.max(x_train, 0)
x_train = 2*(x_train - xmin)/(xmax - xmin) -1

# 测试集复用训练集参数做归一化
x_test = x_train0[n_train:n_train+30]
x_test = 2*(x_test - xmin)/(xmax - xmin) -1

2. 应对小数据集过拟合

训练集仅70样本,测试集30样本,数据量过小是核心问题之一:

  • 采用交叉验证替代单一划分:比如5折交叉验证,更准确评估模型真实性能
  • 尝试结构化数据增强:给特征添加微小噪声、构造特征组合(如多项式特征)、样本权重调整等
  • 简化模型结构:当前3层隐藏层(50+25+50神经元)对小数据集过于复杂,可尝试减少为1-2层隐藏层,或降低神经元数量(如20+10)

3. 添加正则化约束

当前模型无任何正则化,直接针对过拟合添加:

  • Dropout层:在隐藏层激活后添加,随机丢弃部分神经元减少依赖
  • L2正则化:给Dense层添加权重约束,避免权重过大
  • 早停机制:监控验证集损失,停止无效迭代,保留最优模型
from tensorflow.keras.callbacks import EarlyStopping
from tensorflow.keras.regularizers import l2

def my_model(x):
    x = Dense(50, kernel_regularizer=l2(0.001))(x)
    x = LeakyReLU()(x)
    x = tf.keras.layers.Dropout(0.2)(x)
    
    x = Dense(25, kernel_regularizer=l2(0.001))(x)
    x = LeakyReLU()(x)
    x = tf.keras.layers.Dropout(0.2)(x)
    
    x = Dense(5, activation='softmax', name='output1')(x)
    return x

# 早停回调
early_stop = EarlyStopping(monitor='val_loss', patience=30, restore_best_weights=True)
history = model.fit(..., callbacks=[early_stop])

4. 优化训练配置

当前优化器设置不合理:

  • SGD学习率0.001过低且无动量,收敛速度慢且易陷入局部最优,可调整为SGD(lr=0.01, momentum=0.9)
  • 更换为Adam优化器:自适应学习率更适合小数据集训练
optim = tf.keras.optimizers.Adam(learning_rate=0.001)

5. 启用Batch Normalization

导入了BatchNormalization但未使用,在隐藏层激活前添加可稳定训练分布,提升泛化能力:

def my_model(x):
    x = Dense(50)(x)
    x = BatchNormalization()(x)
    x = LeakyReLU()(x)
    
    x = Dense(25)(x)
    x = BatchNormalization()(x)
    x = LeakyReLU()(x)
    
    x = Dense(5, activation='softmax', name='output1')(x)
    return x

内容的提问来源于stack exchange,提问作者gingerorange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 09:44:52