为何多分类神经网络训练时损失值随epoch递增?
我在处理一个7分类的森林覆盖类型任务,采用TensorFlow和Keras框架搭建模型,最后层设置7个节点并使用softmax激活函数,目标标签已转换为one-hot编码,损失函数选用categorical_crossentropy。但训练过程中损失值随每个epoch不断递增,损失变化图表如下:
使用的代码如下:
from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split(X, y, stratify=y, test_size=0.7, shuffle=True, random_state=42) y_train = tf.one_hot(y_train,7,) y_test = tf.one_hot(y_test,7,) # Define the model architecture model = tf.keras.Sequential([ tf.keras.layers.Dense(10, activation='relu'), tf.keras.layers.Dense(128, activation='relu'), tf.keras.layers.Dense(7, activation='softmax') ]) # Compile the model model.compile(tf.keras.optimizers.Adam(learning_rate=0.001), loss='categorical_crossentropy', metrics=['accuracy']) # Train the model history = model.fit(X_train, y_train, epochs=10, batch_size=128) # Evaluate the model on the test data test_loss, test_acc = model.evaluate(X_test, y_test)
以下是几个可能的问题点和解决办法:
训练集占比过低:当前
test_size=0.7,训练集仅占总数据的30%,数据量太小会导致模型无法学到有效特征,训练稳定性极差。建议调整划分比例,比如将test_size改为0.2或0.3,让训练集占大部分数据。特征未做归一化处理:森林覆盖数据集的特征数值跨度可能极大(比如海拔、距离水源的距离等),未归一化的大数值特征会干扰梯度更新方向,导致模型训练失控、损失上升。用标准化工具对输入特征做预处理:
from sklearn.preprocessing import StandardScaler scaler = StandardScaler() X_train = scaler.fit_transform(X_train) X_test = scaler.transform(X_test)标签索引不匹配:如果原始标签的取值是1-7而非0-6,
tf.one_hot会生成8维向量(索引从0开始),但模型最后层只有7个节点,这会导致标签与输出维度不匹配,损失计算异常。先将标签转换为0-6的范围:y_train = y_train - 1 y_test = y_test - 1 y_train = tf.one_hot(y_train, 7) y_test = tf.one_hot(y_test, 7)网络结构与学习率不合理:第一层仅10个神经元可能不足以提取基础特征;同时,若特征未归一化,0.001的学习率可能偏大,导致梯度震荡或爆炸。可以尝试将第一层神经元数量调整为64,或者把学习率降低到0.0001。
数据类型与维度检查:确认
y_train转one-hot后的维度是(样本数,7),与模型输出维度一致;同时将输入特征转换为float32类型,避免计算精度问题:X_train = X_train.astype('float32') X_test = X_test.astype('float32')
内容的提问来源于stack exchange,提问作者suba325

