线性回归框架下,如何用OLS估计方差?兼论SSR/(N-K)是否为LS估计量
Great question—this is a common point of confusion when moving from estimating coefficients to estimating the error variance, especially if you’re coming from an MLE background. Let’s break this down clearly.
First, remember that OLS is primarily about estimating the coefficient vector $\beta$ by minimizing the residual sum of squares (SSR):
$$\hat{\beta} = \arg\min_{\beta} \sum_{i=1}^N (y_i - X_i\beta)^2 = \arg\min_{\beta} SSR(\beta)$$
This gives us the famous closed-form solution $\hat{\beta} = (X'X)^{-1}X'y$. But this only solves for $\beta$—we still need an estimate of the error variance $\sigma^2 = \text{Var}(\varepsilon)$, where $\varepsilon = y - X\beta$ is the true unobserved error.
Let’s start with what we know:
- The residual vector is $\hat{\varepsilon} = y - X\hat{\beta}$, so $SSR = \hat{\varepsilon}'\hat{\varepsilon}$.
- If we used maximum likelihood estimation (MLE) for $\sigma^2$, we’d take the log-likelihood of the normal linear model, take the derivative with respect to $\sigma^2$, and solve for the MLE: $\hat{\sigma}^2_{MLE} = \frac{SSR}{N}$. But this estimator is biased—its expected value is $\frac{N-K}{N}\sigma^2$, not $\sigma^2$.
The estimator $\frac{SSR}{N-K}$ is an adjusted, unbiased estimator of $\sigma^2$. Here’s why the adjustment matters:
When we estimate $\beta$, we’re using $K$ degrees of freedom (one for each coefficient in $\beta$, including the intercept if you have one). The residuals $\hat{\varepsilon}$ aren’t independent of each other—they satisfy $X'\hat{\varepsilon} = 0$, which imposes $K$ linear constraints. This means only $N-K$ of the residuals are "free" to vary.
Mathematically, we can show that:
$$E[SSR] = E[\hat{\varepsilon}'\hat{\varepsilon}] = (N-K)\sigma^2$$
Dividing both sides by $N-K$ gives us an estimator whose expected value is exactly $\sigma^2$—it’s unbiased.
Short answer: No, not in the strict sense.
OLS is defined by the criterion of minimizing the residual sum of squares to estimate $\beta$. The estimator $\frac{SSR}{N-K}$ isn’t derived from minimizing SSR directly—it’s a post-hoc adjustment to the biased MLE (or a moment-based estimator) to achieve unbiasedness. We adopt it because of its strong statistical properties (unbiasedness, consistency, efficiency under normality), not because it’s part of the core OLS coefficient estimation step.
To put it another way: OLS gives us $\hat{\beta}$ and the SSR, but the choice to divide SSR by $N-K$ instead of $N$ is driven by our desire for an unbiased estimator of $\sigma^2$, not by the least squares minimization itself.
- OLS focuses on minimizing SSR to get $\hat{\beta}$; the error variance estimate is a separate, secondary step.
- The MLE of $\sigma^2$ is $\frac{SSR}{N}$, but it’s biased (though asymptotically unbiased).
- $\frac{SSR}{N-K}$ is an unbiased estimator of $\sigma^2$, adopted for its statistical validity, not as a direct output of the least squares minimization process.
内容的提问来源于stack exchange,提问作者Mario Migliaccio

