You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何神经网络权重需初始化为1/√隐藏节点数量?

Why Initialize Neural Network Weights to $1/\sqrt{n}$ Instead of [-1, 1]?

Awesome question—this is one of those small but critical details that can make or break how smoothly your neural network trains. Let’s break this down in plain terms:

The Problem with [-1, 1] Random Weights

If you just pick weights randomly between -1 and 1, you run into big issues once your network has more than a few layers:

  • Gradient vanishing or exploding: As signals pass through each layer, they get multiplied by these weights. For deep networks:
    • If weights are on the larger side, the weighted sum of inputs becomes huge. For activation functions like sigmoid or tanh, this pushes outputs to the extreme ends (near 0 or 1 for sigmoid), where the gradient is practically zero. When gradients are that small, the network can’t learn anything—this is gradient vanishing.
    • If weights are too small, signals get diluted layer after layer, resulting in tiny outputs that don’t carry enough information to update the weights effectively.
  • Unstable training: Without controlling the weight scale, each layer’s output variance swings wildly from one layer to the next. This makes your loss function jump around erratically during training, slowing down convergence or stopping it entirely.

Why $1/\sqrt{n}$ Fixes This

This initialization strategy (closely related to Xavier initialization for sigmoid/tanh) is all about keeping signal variance consistent across layers. Here’s the gist:

  • Let $n$ be the number of nodes in the next hidden layer (as your book notes). By setting weights in the range $[-1/\sqrt{n}, 1/\sqrt{n}]$, we ensure that the variance of the weighted sum of inputs to each node stays roughly equal to the variance of the inputs themselves.
  • This keeps activation outputs in the "sweet spot" of the activation function—away from saturated regions—so gradients stay large enough for meaningful learning. Training becomes stable, and your network converges much faster than with unconstrained [-1,1] weights.

In short: this trick prevents your network from "choking" on signals or gradients early in training, making the entire learning process far more reliable.

内容的提问来源于stack exchange,提问作者Ethernetz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:37:06