NumPy实现三层神经网络权重偏置导数计算与更新排错
问题根因
维度不匹配问题来自反向传播阶段矩阵乘法顺序、转置操作逻辑混乱,没有和前向传播的张量形状规则对齐。先明确当前网络结构下各参数的基准形状(输入x形状为(1,785),MNIST为10分类,设隐藏层神经元数为h),先确认参数初始化形状符合以下规则:
- 输入层到隐藏层权重
weights_input_layer_to_hidden_layer:形状(785, h) - 隐藏层偏置
bias_on_hidden_layer:形状(h, 1) - 隐藏层到输出层权重
weights_hidden_layer_to_output_layer:形状(h, 10) - 输出层偏置
bias_on_output_layer:形状(10, 1)
修正方案
反向传播的核心规则是:每层误差项形状必须和该层未激活值形状完全一致,权重/偏置梯度形状必须和对应参数形状完全一致。
修正后的反向传播代码
def backpropagation(self, x, y): predicted_value = self.forward_propagation(x) # 去掉predicted_value的多余转置,保证损失导数形状和输出层一致为(10,1) cost_value_derivative = self.loss_function( predicted_value, self.expected_value(y), derivative=True ) print(f"{'-*-'*15} PREDICTION {'-*-'*15}") print(f"Predicted Value: {np.argmax(predicted_value)}") print(f"Actual Value: {y}") print(f"{'-*-'*15}{'-*-'*19}") # 计算输出层误差项,形状(10,1),和输出层未激活值形状一致 delta2 = cost_value_derivative * self.sigmoid( self.output_layer_without_activity, derivative=True ) # W2梯度形状(h,10),和隐藏层到输出层权重形状完全匹配 derivative_W2 = delta2.dot(self.hidden_layer.T).T # b2梯度形状(10,1),和输出层偏置形状完全匹配 derivative_b2 = np.sum(delta2, axis=1, keepdims=True) print(f"Derivative_W2: {derivative_W2.shape}, weights_hidden_layer_to_output_layer: {self.weights_hidden_layer_to_output_layer.shape}") assert derivative_W2.shape == self.weights_hidden_layer_to_output_layer.shape print(f"Derivative_b2: {derivative_b2.shape}, bias_on_output_layer: {self.bias_on_output_layer.shape}") assert derivative_b2.shape == self.bias_on_output_layer.shape # 计算隐藏层误差项,形状(h,1),和隐藏层未激活值形状一致 delta1 = self.weights_hidden_layer_to_output_layer.dot(delta2) * self.sigmoid( self.hidden_layer_without_activity, derivative=True ) # W1梯度形状(785,h),和输入层到隐藏层权重形状完全匹配 derivative_W1 = delta1.dot(x).T # b1梯度形状(h,1),和隐藏层偏置形状完全匹配 derivative_b1 = np.sum(delta1, axis=1, keepdims=True) print(f"Derivative_b1: {derivative_b1.shape}, bias_on_hidden_layer: {self.bias_on_hidden_layer.shape}") assert derivative_b1.shape == self.bias_on_hidden_layer.shape print(f"Derivative_W1: {derivative_W1.shape}, weights_input_layer_to_hidden_layer: {self.weights_input_layer_to_hidden_layer.shape}") assert derivative_W1.shape == self.weights_input_layer_to_hidden_layer.shape return derivative_W2, derivative_b2, derivative_W1, derivative_b1
梯度更新注意事项
- 修正后梯度形状和参数完全一致,直接用原有更新逻辑即可,不会触发广播错误:
self.weights_hidden_layer_to_output_layer -= learning_rate * derivative_W2 self.bias_on_output_layer -= learning_rate * derivative_b2 self.weights_input_layer_to_hidden_layer -= learning_rate * derivative_W1 self.bias_on_hidden_layer -= learning_rate * derivative_b1
- 原有更新语句中变量名
weights_on_hidden_layer_to_output_layer和定义的参数名不一致,注意拼写统一,避免触发变量不存在错误。 - 后续如果改成批量样本训练,只需要在计算偏置梯度时保留批量维度求和逻辑即可,权重梯度计算逻辑不需要改动。
内容的提问来源于stack exchange,提问作者Estevan Cardoso
相关产品推荐
相关产品推荐

