如何在Scikit-learn中获取逻辑回归模型的对数似然值?
Hey there! This is a common point of confusion since scikit-learn doesn’t expose the log-likelihood directly, but we can easily compute it using the tools we already have—let’s break this down step by step.
First: Clarify the Relationship Between log_loss and Log-Likelihood
As you noted, scikit-learn’s log_loss is defined as the negative average log-likelihood of the true labels given the model’s predictions. To translate this to the log-likelihood you need:
log_loss(...)= Positive value (average negative log-likelihood)- Average log-likelihood =
-log_loss(...) - Total log-likelihood (sum over all samples, which is required for likelihood ratio tests) =
-log_loss(...) * number_of_samples
Method 1: Use log_loss to Compute Total Log-Likelihood
Here’s how to calculate the total log-likelihood for your training data (since likelihood ratio tests rely on the data the model was fit to):
from sklearn.linear_model import LogisticRegression from sklearn.metrics import log_loss # Fit your model lr = LogisticRegression() lr.fit(X_train, y_train) # Get predicted probabilities for the training set y_train_prob = lr.predict_proba(X_train) # Calculate average negative log-likelihood avg_neg_log_likelihood = log_loss(y_train, y_train_prob) # Convert to total log-likelihood total_log_likelihood = -avg_neg_log_likelihood * len(y_train)
Method 2: Manual Calculation (More Transparent)
If you want to avoid relying on log_loss and see exactly how the log-likelihood is computed, you can calculate it directly per sample:
import numpy as np # Get predicted probabilities for the positive class y_train_prob_pos = lr.predict_proba(X_train)[:, 1] # Calculate log-likelihood for each sample sample_log_likelihoods = (y_train * np.log(y_train_prob_pos)) + ((1 - y_train) * np.log(1 - y_train_prob_pos)) # Sum to get total log-likelihood total_log_likelihood = sample_log_likelihoods.sum()
This will give you the exact same result as Method 1—just a more explicit way to understand the math behind the metric.
Why Doesn’t Scikit-Learn Expose This Directly?
Scikit-learn’s API prioritizes predictive performance and general usability over statistical inference tools (like log-likelihood, p-values, etc.). If you regularly need inference-focused metrics, you might want to check out statsmodels' Logit class, which directly returns log-likelihood values and other statistical outputs. But for scikit-learn, the above methods work perfectly for your use case.
Quick Tip for Likelihood Ratio Tests
When running a likelihood ratio test:
- Compute the total log-likelihood for your full model (
LL_full) - Compute the total log-likelihood for a null/reduced model (e.g., a model with only an intercept, or fewer features) (
LL_null) - Calculate the test statistic:
LR = -2 * (LL_null - LL_full) - Compare this statistic to a chi-squared distribution with degrees of freedom equal to the difference in the number of parameters between the two models.
内容的提问来源于stack exchange,提问作者Mattia Paterna

