You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NumPy实现三层神经网络权重偏置导数计算与更新排错

问题根因

维度不匹配问题来自反向传播阶段矩阵乘法顺序、转置操作逻辑混乱,没有和前向传播的张量形状规则对齐。先明确当前网络结构下各参数的基准形状(输入x形状为(1,785),MNIST为10分类,设隐藏层神经元数为h),先确认参数初始化形状符合以下规则:

  • 输入层到隐藏层权重weights_input_layer_to_hidden_layer:形状(785, h)
  • 隐藏层偏置bias_on_hidden_layer:形状(h, 1)
  • 隐藏层到输出层权重weights_hidden_layer_to_output_layer:形状(h, 10)
  • 输出层偏置bias_on_output_layer:形状(10, 1)
修正方案

反向传播的核心规则是:每层误差项形状必须和该层未激活值形状完全一致,权重/偏置梯度形状必须和对应参数形状完全一致。

修正后的反向传播代码

def backpropagation(self, x, y):
    predicted_value = self.forward_propagation(x)
    # 去掉predicted_value的多余转置,保证损失导数形状和输出层一致为(10,1)
    cost_value_derivative = self.loss_function(
            predicted_value, self.expected_value(y), derivative=True
        )

    print(f"{'-*-'*15} PREDICTION {'-*-'*15}")
    print(f"Predicted Value: {np.argmax(predicted_value)}")
    print(f"Actual Value: {y}")
    print(f"{'-*-'*15}{'-*-'*19}")

    # 计算输出层误差项,形状(10,1),和输出层未激活值形状一致
    delta2 = cost_value_derivative * self.sigmoid(
        self.output_layer_without_activity, derivative=True
    )
    # W2梯度形状(h,10),和隐藏层到输出层权重形状完全匹配
    derivative_W2 = delta2.dot(self.hidden_layer.T).T
    # b2梯度形状(10,1),和输出层偏置形状完全匹配
    derivative_b2 = np.sum(delta2, axis=1, keepdims=True)

    print(f"Derivative_W2: {derivative_W2.shape}, weights_hidden_layer_to_output_layer: {self.weights_hidden_layer_to_output_layer.shape}")
    assert derivative_W2.shape == self.weights_hidden_layer_to_output_layer.shape
    print(f"Derivative_b2: {derivative_b2.shape}, bias_on_output_layer: {self.bias_on_output_layer.shape}")
    assert derivative_b2.shape == self.bias_on_output_layer.shape

    # 计算隐藏层误差项,形状(h,1),和隐藏层未激活值形状一致
    delta1 = self.weights_hidden_layer_to_output_layer.dot(delta2) * self.sigmoid(
        self.hidden_layer_without_activity, derivative=True
    )
    # W1梯度形状(785,h),和输入层到隐藏层权重形状完全匹配
    derivative_W1 = delta1.dot(x).T
    # b1梯度形状(h,1),和隐藏层偏置形状完全匹配
    derivative_b1 = np.sum(delta1, axis=1, keepdims=True)

    print(f"Derivative_b1: {derivative_b1.shape}, bias_on_hidden_layer: {self.bias_on_hidden_layer.shape}")
    assert derivative_b1.shape == self.bias_on_hidden_layer.shape
    print(f"Derivative_W1: {derivative_W1.shape}, weights_input_layer_to_hidden_layer: {self.weights_input_layer_to_hidden_layer.shape}")
    assert derivative_W1.shape == self.weights_input_layer_to_hidden_layer.shape

    return derivative_W2, derivative_b2, derivative_W1, derivative_b1

梯度更新注意事项

  1. 修正后梯度形状和参数完全一致,直接用原有更新逻辑即可,不会触发广播错误:
self.weights_hidden_layer_to_output_layer -= learning_rate * derivative_W2
self.bias_on_output_layer -= learning_rate * derivative_b2
self.weights_input_layer_to_hidden_layer -= learning_rate * derivative_W1
self.bias_on_hidden_layer -= learning_rate * derivative_b1
  1. 原有更新语句中变量名weights_on_hidden_layer_to_output_layer和定义的参数名不一致,注意拼写统一,避免触发变量不存在错误。
  2. 后续如果改成批量样本训练,只需要在计算偏置梯度时保留批量维度求和逻辑即可,权重梯度计算逻辑不需要改动。

内容的提问来源于stack exchange,提问作者Estevan Cardoso

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 12:45:31