为何sklearn与statsmodels的OLS无截距拟合时R²结果不同?
Great question! This is such a common gotcha when switching between these two libraries—they’re built for slightly different worlds (machine learning vs. traditional statistics), and that’s where the R² split happens. Let’s break this down clearly:
Core Cause: Different Definitions of Total Sum of Squares (TSS)
The R² formula boils down to 1 - (Residual Sum of Squares / Total Sum of Squares), but the two libraries disagree on what TSS should be when you omit the intercept:
With an intercept (
fit_intercept=True): Both follow the standard statistical rulebook. TSS is calculated as the sum of squared differences between your target valuesyand their mean (sum((y - y.mean())²)). This is why their R² values match perfectly here.Without an intercept (
fit_intercept=False):- sklearn shifts to a machine learning-focused TSS definition: it uses the sum of squared target values themselves (
sum(y²)), treating 0 as the baseline instead of the mean ofy. This makes sense for ML scenarios where users might intentionally center data at the origin (e.g., after feature scaling that removes the mean). - statsmodels sticks strictly to textbook statistics: it continues to use the mean-based TSS (
sum((y - y.mean())²)), regardless of whether an intercept is included. This aligns with how R² is taught to measure total variation around the dataset’s mean.
- sklearn shifts to a machine learning-focused TSS definition: it uses the sum of squared target values themselves (
A Quick Code Example to Prove It
Let’s run a tiny experiment to see the difference in action:
import numpy as np from sklearn.linear_model import LinearRegression import statsmodels.api as sm # Generate sample data X = np.array([1, 2, 3, 4, 5]).reshape(-1, 1) y = np.array([2, 4, 5, 7, 8]) # Sklearn: No intercept sk_model = LinearRegression(fit_intercept=False) sk_model.fit(X, y) sk_r2 = sk_model.score(X, y) # Statsmodels: No intercept sm_model = sm.OLS(y, X).fit() sm_r2 = sm_model.rsquared print(f"sklearn R² (no intercept): {sk_r2:.4f}") print(f"statsmodels R² (no intercept): {sm_r2:.4f}")
You’ll see output like:
sklearn R² (no intercept): 0.9953
statsmodels R² (no intercept): 0.9673
The gap comes directly from the TSS calculation:
- sklearn uses
sum(y²) = 158as its denominator - statsmodels uses
sum((y - y.mean())²) = 22.8as its denominator
How to Make Them Match If Needed
If you want consistent R² values between the two libraries:
- Force sklearn to use the stats-style R²: Calculate it manually instead of using
score():y_pred = sk_model.predict(X) rss = ((y - y_pred) ** 2).sum() tss = ((y - y.mean()) ** 2).sum() sk_r2_stats_style = 1 - (rss / tss) # This will equal sm_model.rsquared - Force statsmodels to use the sklearn-style R²: Calculate it using the sum of squared
yvalues:sm_r2_sk_style = 1 - (sm_model.ssr / (y ** 2).sum()) # This will equal sk_model.score(X, y)
内容的提问来源于stack exchange,提问作者abukaj

