You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何sklearn与statsmodels的OLS无截距拟合时R²结果不同?

Why sklearn and statsmodels OLS R² Differ When Skipping the Intercept

Great question! This is such a common gotcha when switching between these two libraries—they’re built for slightly different worlds (machine learning vs. traditional statistics), and that’s where the R² split happens. Let’s break this down clearly:

Core Cause: Different Definitions of Total Sum of Squares (TSS)

The R² formula boils down to 1 - (Residual Sum of Squares / Total Sum of Squares), but the two libraries disagree on what TSS should be when you omit the intercept:

  • With an intercept (fit_intercept=True): Both follow the standard statistical rulebook. TSS is calculated as the sum of squared differences between your target values y and their mean (sum((y - y.mean())²)). This is why their R² values match perfectly here.

  • Without an intercept (fit_intercept=False):

    • sklearn shifts to a machine learning-focused TSS definition: it uses the sum of squared target values themselves (sum(y²)), treating 0 as the baseline instead of the mean of y. This makes sense for ML scenarios where users might intentionally center data at the origin (e.g., after feature scaling that removes the mean).
    • statsmodels sticks strictly to textbook statistics: it continues to use the mean-based TSS (sum((y - y.mean())²)), regardless of whether an intercept is included. This aligns with how R² is taught to measure total variation around the dataset’s mean.

A Quick Code Example to Prove It

Let’s run a tiny experiment to see the difference in action:

import numpy as np
from sklearn.linear_model import LinearRegression
import statsmodels.api as sm

# Generate sample data
X = np.array([1, 2, 3, 4, 5]).reshape(-1, 1)
y = np.array([2, 4, 5, 7, 8])

# Sklearn: No intercept
sk_model = LinearRegression(fit_intercept=False)
sk_model.fit(X, y)
sk_r2 = sk_model.score(X, y)

# Statsmodels: No intercept
sm_model = sm.OLS(y, X).fit()
sm_r2 = sm_model.rsquared

print(f"sklearn R² (no intercept): {sk_r2:.4f}")
print(f"statsmodels R² (no intercept): {sm_r2:.4f}")

You’ll see output like:

sklearn R² (no intercept): 0.9953
statsmodels R² (no intercept): 0.9673

The gap comes directly from the TSS calculation:

  • sklearn uses sum(y²) = 158 as its denominator
  • statsmodels uses sum((y - y.mean())²) = 22.8 as its denominator

How to Make Them Match If Needed

If you want consistent R² values between the two libraries:

  1. Force sklearn to use the stats-style R²: Calculate it manually instead of using score():
    y_pred = sk_model.predict(X)
    rss = ((y - y_pred) ** 2).sum()
    tss = ((y - y.mean()) ** 2).sum()
    sk_r2_stats_style = 1 - (rss / tss)
    # This will equal sm_model.rsquared
    
  2. Force statsmodels to use the sklearn-style R²: Calculate it using the sum of squared y values:
    sm_r2_sk_style = 1 - (sm_model.ssr / (y ** 2).sum())
    # This will equal sk_model.score(X, y)
    

内容的提问来源于stack exchange,提问作者abukaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:41:55