Dropout应置于Batch Normalization之前吗?对Ng课程中层顺序的疑问
Great question—this is such a common gotcha when mixing dropout and batch normalization, especially with recurrent layers like LSTMs! Let’s break down why the Dropout -> BatchNorm -> Dropout sequence feels off, and why your suspicion about Keras implementation issues is spot-on.
First: The Standard Best Practice for Order
In almost all standard deep learning workflows, Batch Normalization (BN) should come before Dropout, not the other way around. Here’s why:
- BN works by normalizing the input to a layer using the mean and variance calculated from the current mini-batch. This stabilizes training by reducing internal covariate shift.
- If you apply Dropout first, you’re randomly zeroing out neurons before BN does its normalization. This introduces unnecessary noise into the mini-batch statistics BN relies on—each batch will have a different set of active neurons, so the mean/variance BN computes won’t reflect the true underlying distribution of the LSTM’s output. Over time, this weakens BN’s ability to stabilize training.
Why the Ng Sequence Might Be Unconventional (and Risky)
That Dropout -> BatchNorm -> Dropout setup is definitely outside the norm, especially for Keras. Here’s what happens in practice with Keras’s implementation:
- During training: The first Dropout randomly deactivates neurons, then BN normalizes the noisy output. The second Dropout adds another layer of random deactivation.
- During inference: Dropout is turned off entirely, so BN uses the moving average of mean/variance it learned during training. But this moving average was computed on data that had random neuron drops—meaning the inference-time input to BN doesn’t match the distribution it was trained on. This creates a train-test distribution shift, which can make your model perform worse on unseen data than it did during training.
What’s a Better Alternative for LSTM Outputs?
For LSTM layers, here are more reliable approaches:
- Use
LSTM -> BatchNorm -> Dropout: This follows the standard order—BN stabilizes the LSTM’s output distribution first, then Dropout applies regularization without messing up BN’s statistics. - Leverage LSTM’s built-in regularization: Keras’s
LSTMlayer has arecurrent_dropoutparameter that applies dropout directly to the recurrent connections (the "memory" part of the LSTM), which is often more effective than applying dropout to the final output alone. - Skip the double Dropout: Most of the time, a single Dropout layer (after BN) is enough for regularization. Adding two layers increases the risk of underfitting, especially if your dataset isn’t massive.
Could There Be an Edge Case Where Ng’s Setup Makes Sense?
It’s possible this was a deliberate choice for demonstration purposes (e.g., showing how different regularization techniques can be combined) or for a very specific task where the double Dropout provided a benefit. But for general use cases, it’s not a setup we’d recommend.
内容的提问来源于stack exchange,提问作者Aleksandar Jovanovic

