You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scikit-learn计算线性回归的p值与t值:关于SSE自由度计算的疑问

Understanding SSE Division and Degrees of Freedom in Scikit-learn Linear Regression

Great question—you’re totally on the right track by digging into the degrees of freedom here, and your intuition about the n-k-1 formula is correct! Let’s break down why you might see SSE divided by x.shape[0]-x.shape[1] instead, and when each makes sense.

First, Let’s Clarify Key Definitions

  • n: Number of observations (equal to x.shape[0])
  • k: Number of predictor variables (original features, not including the intercept)
  • Residual degrees of freedom (df_resid): For a linear regression model that includes an intercept (the default in Scikit-learn), this is indeed n - k - 1. This accounts for the k predictor coefficients plus the intercept term we estimate—each estimated parameter reduces the degrees of freedom by 1.

Why the Discrepancy with x.shape[0]-x.shape[1]?

The difference comes down to whether your feature matrix x includes an explicit intercept column:

  1. Case 1: x does NOT include an intercept column (Scikit-learn default)
    Scikit-learn’s LinearRegression automatically adds an intercept term during fitting, so your x.shape[1] equals k (the number of predictors). In this scenario:

    • Correct df_resid = n - k - 1 = x.shape[0] - x.shape[1] - 1
    • If you see code dividing SSE by x.shape[0]-x.shape[1], this is missing the extra -1 and would be incorrect for a model with an intercept.
  2. Case 2: x DOES include an explicit intercept column
    If you manually added a column of 1s to x to represent the intercept (bypassing Scikit-learn’s automatic handling), then x.shape[1] equals k + 1 (predictors + intercept). Now:

    • df_resid = n - (k + 1) = x.shape[0] - x.shape[1]
    • This matches the division you saw, and it’s completely correct—because the intercept is now counted as one of the features in x.shape[1].

Quick Example to Illustrate

Suppose you have 100 observations and 3 predictor variables:

  • Without an intercept column: x.shape = (100, 3), df_resid = 100 - 3 - 1 = 96
  • With an intercept column: x.shape = (100, 4), df_resid = 100 - 4 = 96

In both cases, the correct degrees of freedom are the same—just calculated differently based on whether x includes the intercept.

Scikit-learn-Specific Note

If you’re using LinearRegression and want to compute df_resid programmatically, you can use:

n = x.shape[0]
df_resid = n - model.coef_.size - 1  # model.coef_ has k elements, +1 for intercept

This works regardless of whether you manually added an intercept (since Scikit-learn will exclude the intercept from coef_ if you set fit_intercept=False).

内容的提问来源于stack exchange,提问作者Roheet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 09:29:07