You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Numpy和Pandas实现带分类交叉熵的反向传播:代码修复

问题分析与修复方案

1. 反向传播误差计算完全错误

当前run()函数中用np.e**-x处理loss得到误差的逻辑完全不符合交叉熵+softmax的梯度推导。对于softmax输出+类别交叉熵,输出层的误差可以简化为预测值与one-hot目标值的差值,无需对loss做指数运算。

2. 反向传播循环逻辑错误

range(len(self._network)-1,-1)语法错误,缺少步长参数;同时索引处理错误,反向传播需要从输出层(最后一层)开始往前遍历各层。

3. Softmax实现错误

计算时使用axis=0会按列归一化,但样本是按行存储的,应该按样本维度(axis=1)计算,并加上keepdims=True保证维度匹配,否则会导致归一化结果错误。

4. 前向传播调用不一致

代码中同时出现forward和forward_pass两种调用,需统一为同一个方法(假设forward是正确的前向传播方法)。

修复后的关键代码

修复后的run()函数

def run(self, **kwargs):
    epochs = kwargs['epochs']
    # 将类别标签转为one-hot编码(适用于目标值为形状(8,)的情况)
    num_classes = self._network[-1].output.shape[1]
    one_hot_targets = np.eye(num_classes)[self._target_values]
    
    for epoch in range(epochs):
        # 前向传播
        self._network[0].forward(self._inputs)
        for i in range(len(self._network)-1):
            self._network[i+1].forward(self._network[i].output)
        
        # 计算并打印平均loss,监控训练过程
        loss, correct_confidences = self.evaluate()
        avg_loss = np.mean(loss)
        print(f'Epoch {epoch+1}/{epochs}, Average Loss: {avg_loss:.4f}')
        
        # 输出层误差:softmax+交叉熵的简化梯度
        output_error = self._network[-1].output - one_hot_targets
        # 对误差做平均,避免样本数量影响学习率
        output_error /= len(self._inputs)
        
        # 从最后一层开始反向传播
        current_error = output_error
        for i in range(len(self._network)-1, -1, -1):
            current_error = self._network[i].backward(current_error, self._learning_rate)
    
    # 测试阶段前向传播
    self._network[0].forward(self._testing_data)
    for i in range(len(self._network)-1):
        self._network[i+1].forward(self._network[i].output)
    print("The network's testing outputs were:", self._network[-1].output)

修复后的softmax实现

elif self._activation_function.lower() == 'softmax':
    # 按样本维度计算max和sum,keepdims保证维度匹配
    exp_values = np.exp(neuron_output - np.max(neuron_output, axis=1, keepdims=True))
    neuron_output = exp_values / np.sum(exp_values, axis=1, keepdims=True)

优化后的evaluate()函数(可选)

新增平均loss返回,方便训练监控:

def evaluate(self):
    samples = len(self._network[-1].output)
    y_pred_clipped = np.clip(self._network[-1].output, 1e-7, 1-1e-7)

    if len(self._target_values.shape) == 1:
        correct_confidences = y_pred_clipped[range(samples), self._target_values[:samples]]
    elif len(self._target_values.shape) == 2:
        correct_confidences = np.sum(y_pred_clipped * self._target_values[:samples], axis=1)

    loss = -np.log(correct_confidences)
    avg_loss = np.mean(loss)
    return loss, avg_loss, correct_confidences

内容的提问来源于stack exchange,提问作者CobraCoder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 00:39:24