迪基-富勒检验与简单t检验的差异及AR模型t统计量分布疑问
Great question—this is a common point of confusion when moving from cross-sectional regression to time series unit root testing. Let's unpack this clearly:
The two tests serve entirely different purposes and rely on incompatible assumptions:
- Core Goal: A standard t-test checks if a regression coefficient differs from a fixed value (usually 0) in a stationary linear model. The DF test is specialized to detect unit roots (non-stationarity) in time series—specifically, testing whether $\beta_1 = 1$ in the AR(1) model $Y_t = \beta_1 Y_{t-1} + \varepsilon_t$.
- Assumption Violation: The t-test depends on Gauss-Markov assumptions, including that the regressor is stationary and independent of the error term. In the DF scenario, if $\beta_1 = 1$, $Y_t$ is a random walk (non-stationary), so $Y_{t-1}$ is correlated with cumulative past errors—this breaks the t-test's foundational assumptions.
- Critical Values: For t-tests, we use the standard t-distribution. For DF tests, the test statistic follows a non-standard, skewed distribution (derived from Brownian motion) with more negative critical values than standard t-values (you'll see these in dedicated DF test tables).
The key issue is the behavior of non-stationary data under the null hypothesis ($\beta_1 = 1$):
When $\beta_1 = 1$, $Y_t$ becomes a random walk: $Y_t = Y_{t-1} + \varepsilon_t$, which expands to $Y_t = \sum_{i=1}^t \varepsilon_i$. Here, $Y_{t-1}$ is a cumulative sum of past errors, so it's not stationary—its variance grows with time.
When calculating the t-statistic for $\hat{\beta_1}$:
$$t = \frac{\hat{\beta_1} - 1}{s_{\hat{\beta_1}}}$$
Under the null, both the numerator and denominator converge to random variables tied to Brownian motion (continuous-time random walk), not the normal distribution we see in stationary regressions. This non-normal asymptotic behavior is why the t-statistic doesn't follow a standard t-distribution.
Short answer: You can calculate $\hat{\beta_1}$ using the standard OLS formula you provided (note: I fixed the denominator sign below—should be $\sum x_i^2 - n\bar{x}^2$, not the reverse):
$$\hat\beta_1=\frac{\sum_ix_iy_i-n\bar x\bar y}{\sum_ix_i^2-n\bar x^2}$$
OLS will still give you a consistent estimator of $\beta_1$. However:
- The standard error formula for stationary OLS does not apply here. When the series is non-stationary, the variance of $Y_{t-1}$ grows with $t$, so the denominator of $\hat{\beta_1}$ scales with $t^2$, while the numerator scales with $t$. This means the standard error (derived for stationary data) doesn't account for the non-stationary scaling, leading to misleading values.
- Even if you compute the standard error using the stationary formula, the resulting t-statistic won't follow a standard t-distribution. Using standard t-critical values would give incorrect p-values, leading you to misjudge whether a unit root exists.
内容的提问来源于stack exchange,提问作者honeybadger

