You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

实现通用反向传播:全连接层梯度维度疑问求助

Understanding Backpropagation Dimensions for Fully Connected Layers

Let's break down your dimension questions and work through the issues in your code step by step:

Expected Dimensions (When N=1, Single Sample)

  • dy (gradient from next layer): This is the gradient of the loss with respect to your current layer's output (self.y). Since self.y has shape (1,10), dy must match this dimension: (1,10). It captures how much each element of your current layer's output impacts the total loss.
  • dz (derivative of activation function):
    • For element-wise activations (ReLU, Sigmoid, "none"): The derivative is calculated per element of self.z, so dz shares the same shape as self.z: (1,10).
    • For Softmax: The derivative is a Jacobian matrix (each output element's derivative relative to every input element). For a single sample, this becomes a (10,10) matrix.
  • self.d (gradient for current layer's input): This is the gradient of the loss with respect to your current layer's input (self.X). Since self.X is (1,128), self.d must match this shape: (1,128). It tells you how much each input element contributes to the total loss.

Issues in Your Code & Fixes

Let's go through the key problems and adjust your code accordingly:

1. Softmax Derivative Calculation

Your current code treats the entire batch as a single sample when computing the Softmax Jacobian, which only works for N=1. For batches with multiple samples, you need to compute a Jacobian matrix per sample, resulting in a (N,10,10) tensor:

elif self.activator == 'soft-max':
    s = self.y  # Softmax derivative uses the activation output, not pre-activation z!
    dz = np.zeros((s.shape[0], s.shape[1], s.shape[1]))
    for i in range(s.shape[0]):
        si = s[i].reshape(-1,1)
        dz[i] = np.diagflat(si) - np.dot(si, si.T)

2. Gradient Calculation for self.d, self.W_grad, self.b_grad

You're using dy directly in your calculations, but we need to use the gradient of the loss with respect to the pre-activation z (dL/dz). This is computed as dL/dy * dy/dz (element-wise product for element-wise activations, matrix multiplication for Softmax):

For Element-wise Activations (ReLU, Sigmoid, "none"):

dL_dz = dz * dy  # Element-wise product, shape (N,10)
self.d = np.dot(dL_dz, self.W.T)  # Correct shape (N,128)
self.W_grad = np.dot(self.X.T, dL_dz)  # Correct shape (128,10)
self.b_grad = np.sum(dL_dz, axis=0).reshape(1, -1)  # Correct shape (1,10)

For Softmax:

Since dz is a (N,10,10) tensor, use Einstein summation to compute dL/dz for each sample:

dL_dz = np.einsum('ni,nij->nj', dy, dz)  # Shape (N,10)
# Use this dL_dz for self.d, self.W_grad, self.b_grad as above

3. Sigmoid Derivative Correction

Your Sigmoid derivative uses self.z, but the correct formula uses the activation output self.y (since σ(z)*(1-σ(z)) where σ(z) is self.y):

elif self.activator == 'sigmoid':
    dz = self.y * (1 - self.y)

Final Adjusted Backward Function

Here's your function with all fixes applied:

def backward(self, dy):
    if self.activator == 'relu':
        dz = np.zeros_like(self.z)
        dz[self.z > 0] = 1
        dL_dz = dz * dy
    elif self.activator == 'sigmoid':
        dz = self.y * (1 - self.y)
        dL_dz = dz * dy
    elif self.activator == 'soft-max':
        s = self.y
        dz = np.zeros((s.shape[0], s.shape[1], s.shape[1]))
        for i in range(s.shape[0]):
            si = s[i].reshape(-1, 1)
            dz[i] = np.diagflat(si) - np.dot(si, si.T)
        dL_dz = np.einsum('ni,nij->nj', dy, dz)
    elif self.activator == 'none':
        dz = np.ones_like(self.z)
        dL_dz = dz * dy

    self.d = np.dot(dL_dz, self.W.T)
    self.W_grad = np.dot(self.X.T, dL_dz)
    self.b_grad = np.sum(dL_dz, axis=0).reshape(1, -1)

内容的提问来源于stack exchange,提问作者Mark.F

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:02:18