Sklearn Logistic Regression的One-vs-Rest(OvR)模式中概率归一化机制是怎样的?
Great question! Let's break this down clearly:
First, in a pure one-vs-rest (OvR) setup, each binary classifier is trained independently to distinguish one class from all others. Each of these classifiers outputs a probability (via the sigmoid function) that a sample belongs to its target class. These probabilities are indeed independent in theory—there's no inherent reason their sum would equal 1.
But scikit-learn's
LogisticRegressionwithmulti_class='ovr'applies a normalization step to the raw sigmoid outputs when you callpredict_proba. Here's the exact process:- For each sample, collect the raw probability output from every binary OvR classifier (let’s call these values
p₁, p₂, ..., pₖwherekis the total number of classes). - Calculate the sum of all these raw probabilities:
total = p₁ + p₂ + ... + pₖ. - Normalize each raw probability by dividing it by this total:
final_pᵢ = pᵢ / total.
- For each sample, collect the raw probability output from every binary OvR classifier (let’s call these values
This normalization step ensures the final probabilities sum to 1, making the output consistent with the valid probability distribution format you’d get from the multinomial (softmax) setting.
Looking at your sample output:
array([[0.16973178, 0.46755188, 0.36271634], [0.58228627, 0.0928127 , 0.32490103], [0.28241256, 0.51175978, 0.20582766], ..., [0.17922774, 0.71300755, 0.10776471], [0.05888508, 0.24924809, 0.69186683], [0.25808835, 0.68599321, 0.05591844]])
Each row sums to 1 precisely because scikit-learn has applied this normalization to the raw sigmoid outputs from the three binary OvR classifiers.
It’s worth noting that this normalization is a deliberate design choice in scikit-learn’s implementation, not a requirement of the OvR method itself. The goal is to provide a consistent interface across different multi-class strategies, so users can rely on predict_proba returning valid probability distributions no matter which multi_class setting they use.
内容的提问来源于stack exchange,提问作者Hossein

