实现通用反向传播:全连接层梯度维度疑问求助
Let's break down your dimension questions and work through the issues in your code step by step:
Expected Dimensions (When N=1, Single Sample)
- dy (gradient from next layer): This is the gradient of the loss with respect to your current layer's output (
self.y). Sinceself.yhas shape(1,10),dymust match this dimension:(1,10). It captures how much each element of your current layer's output impacts the total loss. - dz (derivative of activation function):
- For element-wise activations (ReLU, Sigmoid, "none"): The derivative is calculated per element of
self.z, sodzshares the same shape asself.z:(1,10). - For Softmax: The derivative is a Jacobian matrix (each output element's derivative relative to every input element). For a single sample, this becomes a
(10,10)matrix.
- For element-wise activations (ReLU, Sigmoid, "none"): The derivative is calculated per element of
- self.d (gradient for current layer's input): This is the gradient of the loss with respect to your current layer's input (
self.X). Sinceself.Xis(1,128),self.dmust match this shape:(1,128). It tells you how much each input element contributes to the total loss.
Issues in Your Code & Fixes
Let's go through the key problems and adjust your code accordingly:
1. Softmax Derivative Calculation
Your current code treats the entire batch as a single sample when computing the Softmax Jacobian, which only works for N=1. For batches with multiple samples, you need to compute a Jacobian matrix per sample, resulting in a (N,10,10) tensor:
elif self.activator == 'soft-max': s = self.y # Softmax derivative uses the activation output, not pre-activation z! dz = np.zeros((s.shape[0], s.shape[1], s.shape[1])) for i in range(s.shape[0]): si = s[i].reshape(-1,1) dz[i] = np.diagflat(si) - np.dot(si, si.T)
2. Gradient Calculation for self.d, self.W_grad, self.b_grad
You're using dy directly in your calculations, but we need to use the gradient of the loss with respect to the pre-activation z (dL/dz). This is computed as dL/dy * dy/dz (element-wise product for element-wise activations, matrix multiplication for Softmax):
For Element-wise Activations (ReLU, Sigmoid, "none"):
dL_dz = dz * dy # Element-wise product, shape (N,10) self.d = np.dot(dL_dz, self.W.T) # Correct shape (N,128) self.W_grad = np.dot(self.X.T, dL_dz) # Correct shape (128,10) self.b_grad = np.sum(dL_dz, axis=0).reshape(1, -1) # Correct shape (1,10)
For Softmax:
Since dz is a (N,10,10) tensor, use Einstein summation to compute dL/dz for each sample:
dL_dz = np.einsum('ni,nij->nj', dy, dz) # Shape (N,10) # Use this dL_dz for self.d, self.W_grad, self.b_grad as above
3. Sigmoid Derivative Correction
Your Sigmoid derivative uses self.z, but the correct formula uses the activation output self.y (since σ(z)*(1-σ(z)) where σ(z) is self.y):
elif self.activator == 'sigmoid': dz = self.y * (1 - self.y)
Final Adjusted Backward Function
Here's your function with all fixes applied:
def backward(self, dy): if self.activator == 'relu': dz = np.zeros_like(self.z) dz[self.z > 0] = 1 dL_dz = dz * dy elif self.activator == 'sigmoid': dz = self.y * (1 - self.y) dL_dz = dz * dy elif self.activator == 'soft-max': s = self.y dz = np.zeros((s.shape[0], s.shape[1], s.shape[1])) for i in range(s.shape[0]): si = s[i].reshape(-1, 1) dz[i] = np.diagflat(si) - np.dot(si, si.T) dL_dz = np.einsum('ni,nij->nj', dy, dz) elif self.activator == 'none': dz = np.ones_like(self.z) dL_dz = dz * dy self.d = np.dot(dL_dz, self.W.T) self.W_grad = np.dot(self.X.T, dL_dz) self.b_grad = np.sum(dL_dz, axis=0).reshape(1, -1)
内容的提问来源于stack exchange,提问作者Mark.F

