基于Numpy和Pandas实现带分类交叉熵的反向传播:代码修复
问题分析与修复方案
1. 反向传播误差计算完全错误
当前run()函数中用np.e**-x处理loss得到误差的逻辑完全不符合交叉熵+softmax的梯度推导。对于softmax输出+类别交叉熵,输出层的误差可以简化为预测值与one-hot目标值的差值,无需对loss做指数运算。
2. 反向传播循环逻辑错误
range(len(self._network)-1,-1)语法错误,缺少步长参数;同时索引处理错误,反向传播需要从输出层(最后一层)开始往前遍历各层。
3. Softmax实现错误
计算时使用axis=0会按列归一化,但样本是按行存储的,应该按样本维度(axis=1)计算,并加上keepdims=True保证维度匹配,否则会导致归一化结果错误。
4. 前向传播调用不一致
代码中同时出现forward和forward_pass两种调用,需统一为同一个方法(假设forward是正确的前向传播方法)。
修复后的关键代码
修复后的run()函数
def run(self, **kwargs): epochs = kwargs['epochs'] # 将类别标签转为one-hot编码(适用于目标值为形状(8,)的情况) num_classes = self._network[-1].output.shape[1] one_hot_targets = np.eye(num_classes)[self._target_values] for epoch in range(epochs): # 前向传播 self._network[0].forward(self._inputs) for i in range(len(self._network)-1): self._network[i+1].forward(self._network[i].output) # 计算并打印平均loss,监控训练过程 loss, correct_confidences = self.evaluate() avg_loss = np.mean(loss) print(f'Epoch {epoch+1}/{epochs}, Average Loss: {avg_loss:.4f}') # 输出层误差:softmax+交叉熵的简化梯度 output_error = self._network[-1].output - one_hot_targets # 对误差做平均,避免样本数量影响学习率 output_error /= len(self._inputs) # 从最后一层开始反向传播 current_error = output_error for i in range(len(self._network)-1, -1, -1): current_error = self._network[i].backward(current_error, self._learning_rate) # 测试阶段前向传播 self._network[0].forward(self._testing_data) for i in range(len(self._network)-1): self._network[i+1].forward(self._network[i].output) print("The network's testing outputs were:", self._network[-1].output)
修复后的softmax实现
elif self._activation_function.lower() == 'softmax': # 按样本维度计算max和sum,keepdims保证维度匹配 exp_values = np.exp(neuron_output - np.max(neuron_output, axis=1, keepdims=True)) neuron_output = exp_values / np.sum(exp_values, axis=1, keepdims=True)
优化后的evaluate()函数(可选)
新增平均loss返回,方便训练监控:
def evaluate(self): samples = len(self._network[-1].output) y_pred_clipped = np.clip(self._network[-1].output, 1e-7, 1-1e-7) if len(self._target_values.shape) == 1: correct_confidences = y_pred_clipped[range(samples), self._target_values[:samples]] elif len(self._target_values.shape) == 2: correct_confidences = np.sum(y_pred_clipped * self._target_values[:samples], axis=1) loss = -np.log(correct_confidences) avg_loss = np.mean(loss) return loss, avg_loss, correct_confidences
内容的提问来源于stack exchange,提问作者CobraCoder
相关产品推荐
相关产品推荐

