Keras中rate=1的Dropout行为不符合预期的问题排查
Ah, I see what's happening here—great question! Let's break this down step by step.
First, let's recap your setup: you set Dropout(1.0), which should mean every neuron from the previous layer is dropped during training, leaving only zeros as input to the final Dense(10, softmax) layer. You'd expect the model to stop learning since there's no signal flowing through, but accuracy keeps going up. Here's why:
The key culprit: the final dense layer's bias terms
Even when the Dropout layer outputs all zeros, the final softmax layer's bias parameters are still being updated. Let's break the math down quickly:
- When the input to the softmax layer is all zeros, the pre-activation values are just the bias terms:
z = 0*W + b = b - The softmax output is computed from these biases, so initially it's a roughly uniform distribution (since biases start near zero)
- During backpropagation, the gradient for the bias terms doesn't depend on the input (unlike the weights). The gradient for bias
b_iisy_pred_i - y_true_i, which is non-zero as long as the prediction doesn't match the label.
So even with all Dropout neurons dropped, the model can still learn to adjust the final layer's biases to push the softmax output toward the correct class. That's why your accuracy keeps rising—those bias terms are being optimized to make better predictions, even without any input from the earlier layers.
Quick test to confirm this
If you want to verify this, modify your final dense layer to exclude biases:
softmax2 = keras.layers.Dense(10, activation='softmax', name='Softmax2', use_bias=False)(dropout)
Now, with no biases to update, the model's accuracy should stay stuck around 10% (random guessing for 10 classes), since the input to softmax is always zeros, leading to a uniform output distribution.
Another way to check: after training, inspect the final layer's biases—you'll see they've shifted away from their initial values, leaning toward the classes in your training data.
Bonus note on Dropout behavior
Just to clarify: Keras' Dropout layer does correctly apply the dropout rate during training (when training=True, which is automatically set by model.fit()). The issue here isn't that Dropout isn't working—it's that the final layer's biases can still learn even with zero input.
内容的提问来源于stack exchange,提问作者Daniel H. Leung

