You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Wide & Deep神经网络可直接使用未归一化数据?

Why TensorFlow Wide & Deep Works Without Feature Normalization (Unlike Standard DNNs)

Great question! This is a super common confusion when switching from vanilla fully-connected DNNs to Wide & Deep models, so let's break down the key reasons step by step:

  • Adaptive Optimizers Are Built Into the Pipeline
    Standard DNNs often rely on SGD or Momentum optimizers, which are not scale-invariant. If you feed in unnormalized features (like age in 0-100 vs capital_gain with values up to 99999), the gradients from the larger-scale features will completely dominate parameter updates, causing training to oscillate or diverge.
    Wide & Deep, though, uses FTRL (Follow The Regularized Leader) for the Wide linear component and Adam for the Deep DNN component. Both are adaptive optimizers:

    • FTRL automatically adjusts learning rates per feature and applies strong L1 regularization, which helps it handle varying feature scales and even sparse, high-magnitude features like capital_gain.
    • Adam tracks first and second moments of gradients, scaling each parameter's update based on its own gradient history. This means it can compensate for large differences in feature scales without manual normalization.
  • Dual Model Architecture Splits the Workload
    The Wide component is designed for "memorization"—it captures simple, linear relationships between features and the target. Even with unnormalized features, this linear layer (paired with FTRL) can quickly learn meaningful weights without getting thrown off by scale differences.
    The Deep component handles "generalization," but since the Wide part already learns basic patterns, the Deep layer doesn't have to fight against extreme gradient imbalances right out the gate. This makes the overall training process more stable than a standalone DNN that has to learn everything from scratch with unnormalized data.

  • Practical Example Context
    The official Wide & Deep examples (like the adult income prediction) often skip normalization to keep the code simple and focus on demonstrating the model's core functionality. That said, normalization still helps! Even with adaptive optimizers, scaling features to zero-mean unit-variance will typically speed up convergence and improve final performance. The example just shows the model can work without it, not that it's the best practice.


内容的提问来源于stack exchange,提问作者teddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:12:57