决策树中为何选用交叉熵损失而非0/1损失?
Great question—this is such a common "wait, why not just use the most intuitive loss?" moment when you're getting deep into classification models. Let's break this down clearly:
The Core Problem with 0/1 Loss
First, a quick recap: 0/1 loss gives you a score of 0 if your model's prediction is correct, and 1 if it's wrong. It sounds perfect on paper—after all, we care about whether the model gets the answer right or wrong! But it has three critical flaws that make it nearly useless for training most modern models:
It's non-differentiable almost everywhere
Most of our go-to training methods (like gradient descent) rely on calculating gradients to update model parameters. With 0/1 loss, the gradient is 0 at every point except the exact threshold where the prediction flips from correct to incorrect. That means your model has no way to "learn" how to adjust its weights—you can't get a signal like "your prediction was close, tweak this parameter a little" or "you're way off, make a big change." It's like trying to navigate a dark room with only a light that turns on once you're already at the door.It doesn't penalize "how wrong" you are
Suppose you're classifying cats vs. dogs. If your model outputs a 0.4 probability for "cat" (wrong, since the true label is dog) vs. a 0.01 probability for "cat" (also wrong), 0/1 loss treats both cases exactly the same—you get a 1 either way. But the second prediction is way more confident in the wrong answer, and your model should get a much stronger signal to fix that. 0/1 loss ignores this nuance entirely.It's brittle to noise and uncertainty
Real-world data isn't perfect—labels get misassigned, and some samples are genuinely ambiguous. 0/1 loss forces your model to take a hard stance every time, even when the data is messy. This can lead to overfitting to noisy labels, since the model will try to squeeze every sample into a correct/incorrect box instead of learning a robust probability distribution.
Why Cross-Entropy (or Mutual Information-Based Losses) Work So Well
Cross-entropy (and related losses tied to mutual information) fixes all these issues by framing the problem as matching probability distributions instead of just counting correct/incorrect guesses:
It's fully differentiable
You get a smooth, continuous gradient signal across all possible predictions. This lets gradient descent do its job—your model can iteratively adjust weights to reduce the loss, even when it's close to getting the answer right.It penalizes confident wrong answers heavily
Going back to the cat/dog example: a 0.01 probability for the correct label will result in a much higher cross-entropy loss than a 0.4 probability. This gives the model clear feedback on how far off it is, pushing it to become more confident in the right answers and less confident in the wrong ones.It embraces uncertainty
Cross-entropy works with the model's predicted probabilities, not just hard classifications. This means it can handle ambiguous samples gracefully—instead of forcing a wrong hard prediction, the model can learn to output a more uncertain distribution, which is often more realistic for messy real-world data.
At the end of the day, 0/1 loss is great for evaluating how well your final model performs (since it directly measures accuracy), but it's terrible for training because it doesn't give your model the incremental feedback it needs to learn.
内容的提问来源于stack exchange,提问作者Bratt Swan

