You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在神经网络中实现Softmax激活、交叉熵损失及其导数遇阻

嘿,我明白你现在的困境——把回归MLP改成支持多分类,卡在Softmax和交叉熵的实现上了对吧?这俩其实搭配起来有个超实用的简化技巧,能让导数计算轻松不少,我一步步给你捋清楚:

一、先搞定Softmax激活函数

Softmax的作用是把输出层的原始得分(logits)转换成每个类别的概率,公式是:

$Softmax(z_i) = \frac{e{z_i}}{\sum_{j=1}C e^{z_j}}$,其中C是类别数

但直接计算容易出现指数爆炸(比如z_i很大时,e^z_i会溢出),所以一定要做数值稳定处理:先对每个logits减去该批次的最大值,再计算指数和归一化。

代码实现(Python):

import numpy as np

def softmax(z):
    # 数值稳定:减去每个样本的最大logit
    z_stable = z - np.max(z, axis=1, keepdims=True)
    exp_z = np.exp(z_stable)
    return exp_z / np.sum(exp_z, axis=1, keepdims=True)

单独的Softmax导数其实有点复杂,但和交叉熵损失结合后会大幅简化,这也是为什么我们通常把它们放在一起处理的原因。

二、交叉熵损失(结合Softmax)的实现

交叉熵损失用于衡量预测概率和真实标签(one-hot编码)之间的差异,公式是:

$L = -\frac{1}{N} \sum_{i=1}^N \sum_{j=1}^C y_{ij} \log(\hat{y}_{ij})$,其中$\hat{y}$是Softmax的输出,y是真实标签的one-hot矩阵

但直接先算Softmax再算log可能出现数值下溢(比如$\hat{y}_{ij}$接近0时,log会变成-inf),所以更稳妥的方式是合并计算,利用$\log(\frac{e^{z_i}}{\sum e^{z_j}}) = z_i - \log(\sum e^{z_j})$,这样可以避免先算Softmax再取log的问题。

代码实现(支持one-hot标签):

def softmax_cross_entropy(logits, y_true):
    # logits是输出层的原始得分(未经过Softmax),y_true是one-hot编码的标签
    N = logits.shape[0]
    # 计算每个样本的log(sum(exp(logits))),同样做数值稳定
    log_sum_exp = np.log(np.sum(np.exp(logits - np.max(logits, axis=1, keepdims=True)), axis=1, keepdims=True))
    # 合并后的交叉熵计算
    cross_entropy = np.mean(np.max(logits, axis=1, keepdims=True) + log_sum_exp - np.sum(y_true * logits, axis=1, keepdims=True))
    return cross_entropy

三、最关键的:反向传播的导数计算

这部分是你可能最头疼的,但幸运的是,当Softmax和交叉熵结合时,损失对logits(输出层的原始输入)的导数异常简洁:$\frac{\partial L}{\partial z_i} = \hat{y}_i - y_i$,其中$\hat{y}$是Softmax输出,y是真实标签。

为什么会这么简单?因为把交叉熵对Softmax输出求导,再乘以Softmax对logits的导数,化简后就得到这个结果,省去了复杂的矩阵运算。

所以在你的MLP反向传播中,输出层的梯度计算可以这么写:

def backward(output_layer_logits, y_true, hidden_output, weights_output, weights_hidden, learning_rate, activation):
    N = y_true.shape[0]
    # 1. 计算输出层的梯度:softmax输出 - 真实标签
    y_pred = softmax(output_layer_logits)
    delta_output = (y_pred - y_true) / N  # 除以样本数求平均梯度
    
    # 2. 计算隐藏层的梯度,根据激活函数调整导数
    if activation == 'relu':
        delta_hidden = np.dot(delta_output, weights_output.T) * (hidden_output > 0)
    elif activation == 'sigmoid':
        delta_hidden = np.dot(delta_output, weights_output.T) * hidden_output * (1 - hidden_output)
    elif activation == 'tanh':
        delta_hidden = np.dot(delta_output, weights_output.T) * (1 - hidden_output**2)
    
    # 3. 更新权重
    weights_output -= learning_rate * np.dot(hidden_output.T, delta_output)
    weights_hidden -= learning_rate * np.dot(X.T, delta_hidden)  # X是输入数据
    
    return weights_output, weights_hidden

四、整合到你的MLP中

你只需要在模型里加几个判断:

  • 当任务是多分类时,输出层用Softmax激活,损失用softmax_cross_entropy
  • 反向传播时用上面简化后的梯度公式

举个简单的类结构示例:

class MLP:
    def __init__(self, input_dim, hidden_dim, output_dim, activation='relu', task='regression'):
        self.task = task
        # 初始化权重和偏置
        self.weights_hidden = np.random.randn(input_dim, hidden_dim) * 0.01
        self.bias_hidden = np.zeros((1, hidden_dim))
        self.weights_output = np.random.randn(hidden_dim, output_dim) * 0.01
        self.bias_output = np.zeros((1, output_dim))
        
        # 激活函数映射
        self.act_fn = {
            'sigmoid': lambda x: 1/(1+np.exp(-x)),
            'tanh': np.tanh,
            'relu': lambda x: np.maximum(0, x)
        }
        self.activation = activation
        
        # 任务对应的输出层和损失函数
        if task == 'classification':
            self.output_act = softmax
            self.loss_fn = softmax_cross_entropy
        else:
            self.output_act = lambda x: x  # 回归任务输出无激活
            self.loss_fn = lambda y_pred, y_true: np.mean((y_pred - y_true)**2)
    
    def forward(self, X):
        # 隐藏层前向传播
        hidden_z = np.dot(X, self.weights_hidden) + self.bias_hidden
        hidden_a = self.act_fn[self.activation](hidden_z)
        # 输出层前向传播
        output_z = np.dot(hidden_a, self.weights_output) + self.bias_output
        return output_z, hidden_a
    
    def train_step(self, X, y_true, learning_rate):
        output_z, hidden_a = self.forward(X)
        if self.task == 'classification':
            loss = self.loss_fn(output_z, y_true)
            # 反向传播计算梯度
            y_pred = softmax(output_z)
            delta_output = (y_pred - y_true) / X.shape[0]
            # 计算隐藏层梯度
            if self.activation == 'relu':
                delta_hidden = np.dot(delta_output, self.weights_output.T) * (hidden_a > 0)
            elif self.activation == 'sigmoid':
                delta_hidden = np.dot(delta_output, self.weights_output.T) * hidden_a * (1 - hidden_a)
            elif self.activation == 'tanh':
                delta_hidden = np.dot(delta_output, self.weights_output.T) * (1 - hidden_a**2)
            # 更新权重和偏置
            self.weights_output -= learning_rate * np.dot(hidden_a.T, delta_output)
            self.bias_output -= learning_rate * np.sum(delta_output, axis=0)
            self.weights_hidden -= learning_rate * np.dot(X.T, delta_hidden)
            self.bias_hidden -= learning_rate * np.sum(delta_hidden, axis=0)
        else:
            # 原有回归任务的训练逻辑
            y_pred = output_z
            loss = self.loss_fn(y_pred, y_true)
            delta_output = (y_pred - y_true) / X.shape[0]
            # 隐藏层梯度计算(和分类逻辑一致)
            if self.activation == 'relu':
                delta_hidden = np.dot(delta_output, self.weights_output.T) * (hidden_a > 0)
            elif self.activation == 'sigmoid':
                delta_hidden = np.dot(delta_output, self.weights_output.T) * hidden_a * (1 - hidden_a)
            elif self.activation == 'tanh':
                delta_hidden = np.dot(delta_output, self.weights_output.T) * (1 - hidden_a**2)
            # 更新权重
            self.weights_output -= learning_rate * np.dot(hidden_a.T, delta_output)
            self.bias_output -= learning_rate * np.sum(delta_output, axis=0)
            self.weights_hidden -= learning_rate * np.dot(X.T, delta_hidden)
            self.bias_hidden -= learning_rate * np.sum(delta_hidden, axis=0)
        return loss

内容的提问来源于stack exchange,提问作者KOB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:42:28