TensorFlow中softmax_cross_entropy_with_logits的logit为何是“未缩放对数概率”?
Great question! This term always trips up folks when they first encounter it—let's break it down step by step to make it make sense.
First, let's recall what the softmax function does: it takes your raw logits (like the [1.5, 2.4, 0.7] you mentioned from tf.matmul(X,W) + b) and converts them into valid probabilities that sum to 1. The formula for softmax is:
softmax(z_i) = exp(z_i) / sum(exp(z_j) for all j)
where z_i is a single logit value.
Now, let's flip this around. Suppose we have a probability p_i (the output of softmax). The log probability of that class is just log(p_i). If we substitute the softmax formula into this log probability, we get:
log(p_i) = log(exp(z_i) / sum(exp(z_j))) = z_i - log(sum(exp(z_j)))
Ah, here's the key! The logit z_i is exactly the log probability log(p_i) minus a global scaling term (log(sum(exp(z_j)))). That scaling term is what ensures all the log probabilities add up to the log of 1 (which is 0), making the corresponding probabilities sum to 1.
So when TensorFlow calls logits "unscaled log probability", it's saying:
- Logits are the raw values that would become true log probabilities if you subtracted that global normalization factor.
- They're "unscaled" because they haven't been adjusted by that shared log-sum-exp term to make them proper normalized log probabilities.
Concrete Example
Let's use your logits [1.5, 2.4, 0.7] to see this in action:
- Calculate the sum of exponentials:
exp(1.5) + exp(2.4) + exp(0.7) ≈ 4.48 + 11.02 + 2.01 = 17.51 - Take the log of that sum:
log(17.51) ≈ 2.86 - Subtract this value from each logit to get true log probabilities:
1.5 - 2.86 ≈ -1.36(log of ~0.255)2.4 - 2.86 ≈ -0.46(log of ~0.630)0.7 - 2.86 ≈ -2.16(log of ~0.115)
If you exponentiate these values, they add up to 1 (0.255 + 0.630 + 0.115 = 1), which confirms they're valid probabilities. Your original logits were just these log probabilities before that final scaling step.
To put it simply: when you think of logits as "class scores", you're describing their practical role. When you call them "unscaled log probability", you're describing their mathematical relationship to the actual probabilities we get after softmax.
内容的提问来源于stack exchange,提问作者KiHyun Nam

