为何Wide & Deep神经网络可直接使用未归一化数据?
Great question! This is a super common confusion when switching from vanilla fully-connected DNNs to Wide & Deep models, so let's break down the key reasons step by step:
Adaptive Optimizers Are Built Into the Pipeline
Standard DNNs often rely on SGD or Momentum optimizers, which are not scale-invariant. If you feed in unnormalized features (likeagein 0-100 vscapital_gainwith values up to 99999), the gradients from the larger-scale features will completely dominate parameter updates, causing training to oscillate or diverge.
Wide & Deep, though, uses FTRL (Follow The Regularized Leader) for the Wide linear component and Adam for the Deep DNN component. Both are adaptive optimizers:- FTRL automatically adjusts learning rates per feature and applies strong L1 regularization, which helps it handle varying feature scales and even sparse, high-magnitude features like
capital_gain. - Adam tracks first and second moments of gradients, scaling each parameter's update based on its own gradient history. This means it can compensate for large differences in feature scales without manual normalization.
- FTRL automatically adjusts learning rates per feature and applies strong L1 regularization, which helps it handle varying feature scales and even sparse, high-magnitude features like
Dual Model Architecture Splits the Workload
The Wide component is designed for "memorization"—it captures simple, linear relationships between features and the target. Even with unnormalized features, this linear layer (paired with FTRL) can quickly learn meaningful weights without getting thrown off by scale differences.
The Deep component handles "generalization," but since the Wide part already learns basic patterns, the Deep layer doesn't have to fight against extreme gradient imbalances right out the gate. This makes the overall training process more stable than a standalone DNN that has to learn everything from scratch with unnormalized data.Practical Example Context
The official Wide & Deep examples (like the adult income prediction) often skip normalization to keep the code simple and focus on demonstrating the model's core functionality. That said, normalization still helps! Even with adaptive optimizers, scaling features to zero-mean unit-variance will typically speed up convergence and improve final performance. The example just shows the model can work without it, not that it's the best practice.
内容的提问来源于stack exchange,提问作者teddy

