《Attention is all you need》中Softmax大输入梯度小的疑问及验证
Great question—this is a super common point of confusion when diving into the attention mechanism from Attention Is All You Need. Let’s break this down to connect the paper’s warning with your own observations.
First, let’s ground the paper’s context: it’s specifically talking about dot-product attention, where the input to softmax is the product of query and key vectors. When the dimension of these vectors ($d_k$) is large, the variance of the dot-product scales with $d_k$, leading to very large absolute values in the softmax input.
In this scenario, the softmax output tends to get extremely "sharp": most of the probability mass collapses onto one or a tiny handful of tokens. For example, if one input value is way larger than all others, the corresponding softmax output $p_i$ approaches 1, and all other $p_j$ approach 0. The softmax gradient with respect to input $x_i$ is $p_i(1-p_i)$, which nears 0 when $p_i$ is almost 1. For other inputs $x_j$, the gradient is $-p_i p_j$, which is also tiny because $p_j$ is close to 0. That’s the vanishing gradient problem the paper is highlighting.
Now, your calculation is totally correct—this isn’t a universal rule for softmax gradients! Here’s why your observation holds:
- Softmax gradients depend on relative differences between inputs, not just their absolute sizes. If all input values are large but close to each other (like [100, 101, 99]), the softmax outputs will be roughly uniform (each ~1/3) instead of collapsing.
- When combined with cross-entropy loss, the gradient for the target class incorporates the loss derivative: $-(1-p_i)$. If $p_i$ isn’t near 1 (because other inputs are also large and competitive), $1-p_i$ stays significant, so the overall gradient remains substantial.
To sum up: the paper’s warning applies when dot-product values become disparately large (due to high $d_k$), causing softmax to collapse. But when multiple inputs are large and balanced, the gradient doesn’t vanish—your analysis is spot-on here.
内容的提问来源于stack exchange,提问作者Richard. Zhu

