为何在深度神经网络分类任务中以最小化交叉熵为优化目标?该目标如何帮助模型学习真实概率分布?
Great question—let’s unpack this clearly, since cross entropy’s dominance makes perfect sense once you connect its math to how models actually learn.
First, let’s circle back to what you already understand:
- The entropy of the true distribution
p(X)measures how much uncertainty exists in the true labels. It’s a fixed value for your dataset, sincep(X)is the ground-truth distribution we’re trying to learn. - Cross entropy between
p(X)and your model’s predicted distributionq(X)is calculated asΣ (i=0 to n) p(X_i) × log(q(X_i)). But here’s a key breakdown that ties everything together:Cross Entropy = Entropy(p) + KL Divergence(p || q)
KL Divergence (Kullback-Leibler Divergence) is a non-negative value that quantifies how dissimilar two probability distributions are. It equals zero only when p and q are exactly identical.
Why Minimize Cross Entropy?
Since Entropy(p) is fixed (it’s just a property of your dataset’s true labels), minimizing cross entropy is mathematically identical to minimizing KL Divergence. In plain language: we’re telling the model to make its predicted distribution q(X) as close as possible to the true distribution p(X).
But why not use another loss like Mean Squared Error (MSE) for classification? Cross entropy has a critical advantage in training efficiency:
- When your model makes a bad prediction (e.g., true label is 1, but
q(X)is 0.1), cross entropy’s gradient is large—so the model updates its parameters more aggressively to fix the mistake. - With MSE, gradients shrink as predictions get closer to 0 or 1, which can slow down learning or even cause the model to get stuck in suboptimal states. Cross entropy avoids this by keeping gradients proportional to the prediction error, no matter how extreme the probabilities get.
How This Helps the Model Learn the True Distribution
Every time you update your model’s parameters to reduce cross entropy, you’re nudging q(X) to match p(X) more closely. Here’s how that translates to meaningful learning:
- Maximizing Likelihood: Minimizing cross entropy is mathematically equivalent to maximizing the likelihood of your model predicting the true labels. In statistical terms, we’re finding the model parameters that make the observed data (your true labels) as probable as possible.
- Converging to the True Distribution: As training progresses and cross entropy decreases, KL Divergence approaches zero. When it hits zero,
q(X)=p(X)—meaning your model has learned the exact probability distribution of the true data (or as close as possible given your model’s capacity).
Put simply, cross entropy gives us a direct, efficient way to measure how wrong our model’s probabilistic predictions are, and optimizing it pushes the model to align its understanding of the data with reality.
内容的提问来源于stack exchange,提问作者Chetan Garg

