You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

statsmodels与sklearn逻辑回归结果差异巨大,求原因及解决方法

Why Statsmodels Logit and Sklearn LogisticRegression Results Are Discrepant?

Let's break down the key reasons for the huge differences you're seeing, and how to fix them to get consistent results for inference:

1. Regularization is Enabled by Default in Sklearn

This is the most likely culprit behind your mismatched results:

  • Statsmodels' Logit does not apply any regularization by default—it fits the pure maximum likelihood estimate (MLE) directly.
  • Sklearn's LogisticRegression uses L2 regularization (ridge penalty) out of the box, even when you set C=1e8 (a large C reduces regularization strength, but it doesn't turn it off entirely). With 28 variables, even weak regularization can shrink coefficients drastically, which explains why your sklearn coefficients are near 0 while statsmodels' are much larger.

Fix: Disable Regularization in Sklearn

Use penalty='none' (available in sklearn 0.21+) to match statsmodels' behavior:

clf = LogisticRegression(
    C=1e8, 
    fit_intercept=False, 
    penalty='none',  # Turn off regularization entirely
    solver='lbfgs'   # Align with statsmodels' default BFGS optimizer
)
clf.fit(X_train, y_train)

2. Feature Scaling Sensitivity

Sklearn's regularized models are highly sensitive to feature scales. If your features live on very different scales (e.g., one variable ranges 0-1000, another 0-1), the regularization penalty will disproportionately shrink coefficients for larger-scale features.

Statsmodels' MLE doesn't care about feature scales, so it returns coefficients in the original feature units.

Fix: Use Unscaled Features for Both Models

Make sure you're feeding the exact same raw (unscaled) data to both sm.Logit and sklearn.LogisticRegression. If you previously scaled features for sklearn, revert that to match statsmodels' input.

3. Verify Intercept Handling

You set fit_intercept=False in sklearn, but double-check that:

  • Your X_train does not include a constant term (since you already dropped one category per variable to avoid multicollinearity, this should be true).
  • Statsmodels' Logit is also not using an intercept—your code sm.Logit(y_train, X_train) is correct here, as statsmodels doesn't add an intercept automatically (you'd need sm.add_constant(X_train) if you wanted one).

4. Check Optimizer Convergence

While less likely to cause extreme differences, both libraries use slightly different default optimizers. Statsmodels uses BFGS by default, while sklearn uses 'lbfgs' (a variant) for binary classification. Ensuring you use solver='lbfgs' in sklearn (as in the code above) helps align the optimization process.

What to Expect After Fixing

Once you disable regularization and use the same raw data:

  • Your sklearn coefficients should match (or be extremely close to) statsmodels' coefficients.
  • The sklearn log_loss should equal -statsmodels_log_likelihood / len(y_train). For your numbers: -(-19661)/10348 ≈ 1.9 (which matches your current log_loss), so this checks out mathematically once coefficients align.

内容的提问来源于stack exchange,提问作者Juan C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:54:50