You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何tanh激活函数几乎总是优于sigmoid?相关原理及特性探究

Great question—this is a common point of confusion when first diving into activation functions, so let’s unpack it thoroughly.

Why Centered Outputs Speed Up Learning

First, let’s clarify: when we say this "makes learning easier for the next layer," we’re talking about the entire network’s learning process, but the impact is most directly felt in how the previous layer’s weights are updated during backpropagation. Here’s why centering matters:

  • When a sigmoid activation outputs values centered around 0.5 (all positive, between 0 and 1), the gradients passed back to the previous layer’s weights will often all have the same sign (either all positive or all negative) if the input to the sigmoid is consistently biased in one direction. This means weight updates will tend to move in a single, unbalanced direction—like pushing a ball up a hill only from one side—leading to slower, zig-zagging convergence.
  • Tanh’s output is centered around 0 (ranging from -1 to 1), so the gradients passed back can be positive or negative. This balanced distribution lets weight updates adjust in more nuanced directions, avoiding that one-sided push. It’s like having equal access to nudge the ball from all angles, making it easier to reach optimal weights faster.

In short: centering the activation output ensures the gradients flowing back to the prior layer have a zero mean, which aligns with the idea of whitening input data—stabilizing the input distribution and reducing variance in weight updates, which directly speeds up convergence.

Other Key Advantages of tanh Over Sigmoid

Beyond centering, tanh has a few other traits that make it preferable for most hidden layers:

  • Better gradient flow: The derivative of tanh is 1 - tanh(x)^2, which peaks at 1 when x=0. Compare that to sigmoid’s derivative, which only peaks at 0.25. A larger derivative means gradients are less likely to "vanish" as they propagate back through layers—critical for deeper networks where gradient vanishing can stall learning entirely.
  • Wider output range: Tanh maps inputs to [-1,1], whereas sigmoid maps to [0,1]. This wider range preserves more of the original input’s variance, so less information is lost when passing data between layers. For example, if your input has both positive and negative signals, tanh will retain that contrast, while sigmoid will squash everything into a positive range, blurring those differences.
  • Simpler weight initialization: Since tanh’s output is zero-centered, you don’t have to tweak weight initialization as carefully to compensate for a non-zero mean in activation outputs. This simplifies setting up the network and reduces the chance of initializing weights that lead to slow, stuck convergence.

Note: The one case where sigmoid still shines is often in binary classification output layers, where you need a probability between 0 and 1. For hidden layers though, tanh is almost always the better pick (and these days, ReLU variants are even more common—but that’s a separate discussion!).

内容的提问来源于stack exchange,提问作者Tom Hale

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:23:30