反向传播保留Sigmoid导数精度低,移除后陡升的原因排查
Hey there, let's break down this tricky but super common gotcha you're hitting with backpropagation. The core issue here is that your implementation has a mismatch between how you're calculating gradients (and using activation function derivatives) and either your forward propagation logic or choice of loss function. Let's unpack this step by step:
1. Why Gradient Checking Passes But Accuracy Is Low
First, the fact that gradient checking matches your backprop result means your code is mathematically consistent—but both the analytical gradient and numerical gradient are wrong. Here are the most likely culprits:
a. You're Using the Wrong Value for the Sigmoid Derivative
Andrew Ng's backprop formula uses the derivative of the activation function evaluated at the pre-activation value z, which for Sigmoid is g'(z) = a * (1 - a) (where a is the post-activation output g(z)).
If your forwardpropagation function returns A as the pre-activation z values (instead of the post-activation a values), then A[l] * (1 - A[l]) is not the correct derivative. You'd instead need to compute self.sigmoid(A[l]) * (1 - self.sigmoid(A[l])) to get the right derivative.
This mistake would make your gradients mathematically consistent (hence gradient checking passes) but directionally wrong, preventing the model from learning properly and leading to the 45% accuracy.
b. You're Using Squared Error Loss with Sigmoid Outputs
For multi-class digit recognition, squared error loss paired with Sigmoid outputs leads to gradient vanishing when predictions are confident (close to 0 or 1). The Sigmoid derivative g'(z) becomes tiny in these cases, scaling down the gradient to near-zero—so the model can't update weights effectively.
2. Why Removing the Derivative Improves Accuracy (But Breaks Gradient Checking)
When you remove * A[l] * (1 - A[l]), you're effectively ignoring the activation function's derivative. This means:
- Your gradient is no longer mathematically correct (hence gradient checking fails), but
- You're bypassing the gradient vanishing problem. Without the tiny derivative scaling the gradient, weights can update freely, allowing the model to learn patterns and jump to 90% accuracy. This is a "lucky mistake" that works for a shallow network, but it's not a correct implementation.
Fixes to Get Both Correct Gradients and High Accuracy
Let's walk through the steps to fix this:
Step 1: Verify Forward Propagation Outputs
Check your forwardpropagation function to ensure it returns post-activation values for all layers (except the input layer). Example of a correct implementation:
def forwardpropagation(self, x): A = [x] # Input layer (no activation) current = x for l in range(self.n_layers - 1): z = np.dot(self.theta[l], current) + self.bias[l] current = self.sigmoid(z) # Apply Sigmoid activation A.append(current) return A
If you must keep pre-activation values in A, adjust the derivative calculation in backprop:
# Replace this line # delta = np.dot(self.theta[l], delta) * A[l] * (1 - A[l]) # With this (if A[l] is pre-activation z): a_l = self.sigmoid(A[l]) delta = np.dot(self.theta[l], delta) * a_l * (1 - a_l)
Step 2: Switch to Cross-Entropy Loss for Multi-Class Classification
Squared error is not ideal for multi-class problems. Instead, use cross-entropy loss with either a Softmax or Sigmoid output layer. Cross-entropy loss eliminates gradient vanishing for Sigmoid/Softmax outputs, as the derivative term cancels out the activation's derivative in the gradient calculation.
For Softmax outputs (recommended for 10-class digit recognition), the output layer delta remains A[-1] - y (which you're already using), but the loss function will be more effective at driving learning.
Step 3: Double-Check Matrix Dimensions
You mentioned your theta is transposed from Ng's tutorial. Confirm that np.dot(self.theta[l], delta) is the correct equivalent of Ng's theta[l]^T @ delta[l+1]. Given your theta is transposed, this should be correct, but it's worth verifying with small test inputs to ensure no dimensionality mismatches are causing silent errors.
Step 4: Re-Run Gradient Checking After Fixes
Once you make these changes, gradient checking should still show small differences (around 1e-7 or lower), and your model should achieve accuracy comparable to when you removed the derivative—but now with a mathematically correct implementation.
内容的提问来源于stack exchange,提问作者Gonzalo Solera

