使用Scikit-learn计算线性回归的p值与t值:关于SSE自由度计算的疑问
Great question—you’re totally on the right track by digging into the degrees of freedom here, and your intuition about the n-k-1 formula is correct! Let’s break down why you might see SSE divided by x.shape[0]-x.shape[1] instead, and when each makes sense.
First, Let’s Clarify Key Definitions
n: Number of observations (equal tox.shape[0])k: Number of predictor variables (original features, not including the intercept)- Residual degrees of freedom (df_resid): For a linear regression model that includes an intercept (the default in Scikit-learn), this is indeed
n - k - 1. This accounts for thekpredictor coefficients plus the intercept term we estimate—each estimated parameter reduces the degrees of freedom by 1.
Why the Discrepancy with x.shape[0]-x.shape[1]?
The difference comes down to whether your feature matrix x includes an explicit intercept column:
Case 1:
xdoes NOT include an intercept column (Scikit-learn default)
Scikit-learn’sLinearRegressionautomatically adds an intercept term during fitting, so yourx.shape[1]equalsk(the number of predictors). In this scenario:- Correct df_resid =
n - k - 1=x.shape[0] - x.shape[1] - 1 - If you see code dividing SSE by
x.shape[0]-x.shape[1], this is missing the extra-1and would be incorrect for a model with an intercept.
- Correct df_resid =
Case 2:
xDOES include an explicit intercept column
If you manually added a column of 1s toxto represent the intercept (bypassing Scikit-learn’s automatic handling), thenx.shape[1]equalsk + 1(predictors + intercept). Now:- df_resid =
n - (k + 1)=x.shape[0] - x.shape[1] - This matches the division you saw, and it’s completely correct—because the intercept is now counted as one of the features in
x.shape[1].
- df_resid =
Quick Example to Illustrate
Suppose you have 100 observations and 3 predictor variables:
- Without an intercept column:
x.shape = (100, 3), df_resid =100 - 3 - 1 = 96 - With an intercept column:
x.shape = (100, 4), df_resid =100 - 4 = 96
In both cases, the correct degrees of freedom are the same—just calculated differently based on whether x includes the intercept.
Scikit-learn-Specific Note
If you’re using LinearRegression and want to compute df_resid programmatically, you can use:
n = x.shape[0] df_resid = n - model.coef_.size - 1 # model.coef_ has k elements, +1 for intercept
This works regardless of whether you manually added an intercept (since Scikit-learn will exclude the intercept from coef_ if you set fit_intercept=False).
内容的提问来源于stack exchange,提问作者Roheet

