statsmodels与sklearn逻辑回归结果差异巨大,求原因及解决方法
Let's break down the key reasons for the huge differences you're seeing, and how to fix them to get consistent results for inference:
1. Regularization is Enabled by Default in Sklearn
This is the most likely culprit behind your mismatched results:
- Statsmodels'
Logitdoes not apply any regularization by default—it fits the pure maximum likelihood estimate (MLE) directly. - Sklearn's
LogisticRegressionuses L2 regularization (ridge penalty) out of the box, even when you setC=1e8(a large C reduces regularization strength, but it doesn't turn it off entirely). With 28 variables, even weak regularization can shrink coefficients drastically, which explains why your sklearn coefficients are near 0 while statsmodels' are much larger.
Fix: Disable Regularization in Sklearn
Use penalty='none' (available in sklearn 0.21+) to match statsmodels' behavior:
clf = LogisticRegression( C=1e8, fit_intercept=False, penalty='none', # Turn off regularization entirely solver='lbfgs' # Align with statsmodels' default BFGS optimizer ) clf.fit(X_train, y_train)
2. Feature Scaling Sensitivity
Sklearn's regularized models are highly sensitive to feature scales. If your features live on very different scales (e.g., one variable ranges 0-1000, another 0-1), the regularization penalty will disproportionately shrink coefficients for larger-scale features.
Statsmodels' MLE doesn't care about feature scales, so it returns coefficients in the original feature units.
Fix: Use Unscaled Features for Both Models
Make sure you're feeding the exact same raw (unscaled) data to both sm.Logit and sklearn.LogisticRegression. If you previously scaled features for sklearn, revert that to match statsmodels' input.
3. Verify Intercept Handling
You set fit_intercept=False in sklearn, but double-check that:
- Your
X_traindoes not include a constant term (since you already dropped one category per variable to avoid multicollinearity, this should be true). - Statsmodels'
Logitis also not using an intercept—your codesm.Logit(y_train, X_train)is correct here, as statsmodels doesn't add an intercept automatically (you'd needsm.add_constant(X_train)if you wanted one).
4. Check Optimizer Convergence
While less likely to cause extreme differences, both libraries use slightly different default optimizers. Statsmodels uses BFGS by default, while sklearn uses 'lbfgs' (a variant) for binary classification. Ensuring you use solver='lbfgs' in sklearn (as in the code above) helps align the optimization process.
What to Expect After Fixing
Once you disable regularization and use the same raw data:
- Your sklearn coefficients should match (or be extremely close to) statsmodels' coefficients.
- The sklearn
log_lossshould equal-statsmodels_log_likelihood / len(y_train). For your numbers:-(-19661)/10348 ≈ 1.9(which matches your current log_loss), so this checks out mathematically once coefficients align.
内容的提问来源于stack exchange,提问作者Juan C

