You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scikit-learn中获取逻辑回归模型的对数似然值?

How to Get Log-Likelihood from Scikit-Learn's Logistic Regression for Likelihood Ratio Tests

Hey there! This is a common point of confusion since scikit-learn doesn’t expose the log-likelihood directly, but we can easily compute it using the tools we already have—let’s break this down step by step.

First: Clarify the Relationship Between log_loss and Log-Likelihood

As you noted, scikit-learn’s log_loss is defined as the negative average log-likelihood of the true labels given the model’s predictions. To translate this to the log-likelihood you need:

  • log_loss(...) = Positive value (average negative log-likelihood)
  • Average log-likelihood = -log_loss(...)
  • Total log-likelihood (sum over all samples, which is required for likelihood ratio tests) = -log_loss(...) * number_of_samples

Method 1: Use log_loss to Compute Total Log-Likelihood

Here’s how to calculate the total log-likelihood for your training data (since likelihood ratio tests rely on the data the model was fit to):

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import log_loss

# Fit your model
lr = LogisticRegression()
lr.fit(X_train, y_train)

# Get predicted probabilities for the training set
y_train_prob = lr.predict_proba(X_train)

# Calculate average negative log-likelihood
avg_neg_log_likelihood = log_loss(y_train, y_train_prob)

# Convert to total log-likelihood
total_log_likelihood = -avg_neg_log_likelihood * len(y_train)

Method 2: Manual Calculation (More Transparent)

If you want to avoid relying on log_loss and see exactly how the log-likelihood is computed, you can calculate it directly per sample:

import numpy as np

# Get predicted probabilities for the positive class
y_train_prob_pos = lr.predict_proba(X_train)[:, 1]

# Calculate log-likelihood for each sample
sample_log_likelihoods = (y_train * np.log(y_train_prob_pos)) + ((1 - y_train) * np.log(1 - y_train_prob_pos))

# Sum to get total log-likelihood
total_log_likelihood = sample_log_likelihoods.sum()

This will give you the exact same result as Method 1—just a more explicit way to understand the math behind the metric.

Why Doesn’t Scikit-Learn Expose This Directly?

Scikit-learn’s API prioritizes predictive performance and general usability over statistical inference tools (like log-likelihood, p-values, etc.). If you regularly need inference-focused metrics, you might want to check out statsmodels' Logit class, which directly returns log-likelihood values and other statistical outputs. But for scikit-learn, the above methods work perfectly for your use case.

Quick Tip for Likelihood Ratio Tests

When running a likelihood ratio test:

  1. Compute the total log-likelihood for your full model (LL_full)
  2. Compute the total log-likelihood for a null/reduced model (e.g., a model with only an intercept, or fewer features) (LL_null)
  3. Calculate the test statistic: LR = -2 * (LL_null - LL_full)
  4. Compare this statistic to a chi-squared distribution with degrees of freedom equal to the difference in the number of parameters between the two models.

内容的提问来源于stack exchange,提问作者Mattia Paterna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:53:24