You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现MLP用Softmax+交叉熵时损失爆炸的原因排查

Softmax+交叉熵训练MLP时损失爆炸的原因与解决方法

问题概述

基于NumPy从零实现MLP模型,使用鸢尾花数据集(特征归一化、标签独热编码),当输出层采用Sigmoid激活+MSE损失时,模型训练正常(最终损失约0.123,准确率0.667);但切换为Softmax激活+交叉熵损失时,损失出现爆炸(最终损失达1112.78),准确率仅0.08。

代码实现

激活与损失函数

import numpy as np

# 激活与损失函数
def relu(x):
    return np.maximum(0, x)

def relu_prime(x):
    return np.where(x > 0, 1, 0)

def sigmoid(x):
    return 1 / (1 + np.exp(-x))

def sigmoid_prime(x):
    return sigmoid(x) * (1 - sigmoid(x))

def softmax(x):
    exp = np.exp(x)
    return exp / np.sum(exp, axis=1, keepdims=True)

def softmax_prime(x):
    return softmax(x) * (1 - softmax(x))

def cross_entropy(y, y_hat):
    return -np.sum(y * np.log(y_hat + 1e-8))

def cross_entropy_prime(y, y_hat):
    return y - y_hat

def mse(y, y_hat):
    return np.mean((y - y_hat) ** 2)

def mse_prime(y, y_hat):
    return 2 * (y_hat - y) / y.size

核心类实现

class Layer:
    def __init__(self, n_input, n_neurons, activation=relu, activation_prime=relu):
        self.weights = np.random.randn(n_input, n_neurons)
        self.biases = np.random.randn(n_neurons)
        self.activation = activation
        self.activation_prime = activation_prime

    def forward(self, inputs):
        self.inputs = inputs
        self.z = np.dot(inputs, self.weights) + self.biases
        self.output = self.activation(self.z)
        return self.output

    def backward(self, dvalues):
        self.dz = dvalues * self.activation_prime(self.z)
        self.dinputs = self.dz.dot(self.weights.T)
        self.dweights = self.inputs.T.dot(self.dz)
        self.dbiases = np.sum(self.dz, axis=0)
        return self.dinputs

    def update(self, learning_rate):
        self.weights -= self.dweights * learning_rate
        self.biases -= self.dbiases * learning_rate

class Model:
    def __init__(self):
        self.layers = []

    def add(self, layer):
        self.layers.append(layer)

    def forward(self, inputs):
        for layer in self.layers:
            inputs = layer.forward(inputs)
        return inputs

    def backward(self, dvalues):
        for layer in reversed(self.layers):
            dvalues = layer.backward(dvalues)

    def update(self, learning_rate):
        for layer in self.layers:
            layer.update(learning_rate)

    def predict(self, inputs):
        return self.forward(inputs)

    def evaluate(self, X, Y):
        predictions = self.predict(X)
        return np.mean(np.argmax(predictions, axis=1) == np.argmax(Y, axis=1))

    def compile(self, loss, loss_prime, learning_rate=0.01):
        self.loss = loss
        self.loss_prime = loss_prime
        self.learning_rate = learning_rate

    def fit(self, X, Y, epochs=100):
        loss = []
        for i in range(epochs):
            outputs = self.forward(X)
            loss.append(self.loss(Y, outputs))
            dvalues = self.loss_prime(Y, outputs)
            self.backward(dvalues)
            self.update(self.learning_rate)
            print(f"Epoch {i}: {loss[-1]}")
        return loss

数据集处理

from sklearn.datasets import load_iris

iris = load_iris()
X = iris.data
X = (X - np.min(X, axis=0)) / (np.max(X, axis=0) - np.min(X, axis=0))
Y = iris.target
y = np.zeros((X.shape[0], 3))
y[np.arange(X.shape[0]), Y] = 1
Y = y

两种训练方式及结果

Sigmoid+MSE(训练正常)

model = Model()
model.add(Layer(4, 5))
model.add(Layer(5, 6))
model.add(Layer(6, 3, activation=sigmoid, activation_prime=sigmoid_prime))

model.compile(loss=mse, loss_prime=mse_prime, learning_rate=0.004)
loss = model.fit(X, Y, epochs=20000)
# 最终损失:0.12373022229717626,准确率:0.6666666666666666

Softmax+交叉熵(损失爆炸)

model2 = Model()
model2.add(Layer(4, 5))
model2.add(Layer(5, 6))
model2.add(Layer(6, 3, activation=softmax, activation_prime=softmax_prime))

model2.compile(cross_entropy, cross_entropy_prime, learning_rate=0.00001)
loss = model2.fit(X, Y, epochs=300)
# 最终损失:1112.783115819416,准确率:0.08

问题分析

1. Softmax导数实现错误

Softmax的导数并非element-wise的softmax(x)*(1-softmax(x))(这是Sigmoid的导数形式)。Softmax输出是向量,其导数为雅可比矩阵,但当与交叉熵损失结合时,两者的联合导数可大幅简化,无需单独计算Softmax的雅可比矩阵。当前代码中,输出层反向传播时用错误的Softmax导数乘以损失梯度,导致梯度计算完全错误,权重更新方向混乱,最终损失爆炸。

2. 交叉熵损失未取平均

当前cross_entropy函数对所有样本损失求和而非取平均,导致损失值量级过大(鸢尾花共150个样本,单个样本损失最大约18.4,求和后可达数千),进一步放大梯度错误的影响。

3. 权重初始化与学习率问题

np.random.randn生成的权重标准差为1,对于输入维度较大的层,易导致初始激活值过大,Softmax输出趋近于0或1,交叉熵损失瞬间飙升。加上学习率设置过小(0.00001),即使梯度正确,模型也难以有效更新。

解决方案

1. 修正Softmax+交叉熵的反向传播逻辑

交叉熵损失对Softmax输入z的导数可简化为y_hat - y(或y - y_hat,取决于损失函数定义),无需再乘以Softmax导数。定义一个返回全1的辅助函数作为输出层的激活导数:

def softmax_identity_prime(x):
    return np.ones_like(x)

创建输出层时使用该函数:

model2.add(Layer(6, 3, activation=softmax, activation_prime=softmax_identity_prime))

2. 修正交叉熵损失为平均损失

调整cross_entropy函数为对所有样本取平均,降低损失值量级:

def cross_entropy(y, y_hat):
    return -np.mean(y * np.log(y_hat + 1e-8))

3. 优化权重初始化

采用Xavier初始化,避免初始激活值过大:

# 在Layer类的__init__中替换权重初始化代码
self.weights = np.random.randn(n_input, n_neurons) * np.sqrt(1 / n_input)

4. 调整学习率

将学习率调至合理范围(如0.01),确保模型能有效更新权重。

修正后的训练代码及预期结果

model2 = Model()
model2.add(Layer(4, 5))
model2.add(Layer(5, 6))
# 使用修正后的activation_prime
model2.add(Layer(6, 3, activation=softmax, activation_prime=softmax_identity_prime))

# 使用平均交叉熵损失
model2.compile(cross_entropy, cross_entropy_prime, learning_rate=0.01)
loss = model2.fit(X, Y, epochs=1000)
# 预期结果:损失逐渐下降至0.1以下,准确率接近1.0

内容的提问来源于stack exchange,提问作者Capta1n_n9m0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 02:12:22