如何检验时间序列平稳性?基于高速车速数据集的AR(p)建模问询
Hey there! Since you've already spotted potential seasonality and non-stationarity in your 15-minute interval highway speed dataset (2976 monthly observations), let's walk through the most reliable ways to formally test for stationarity in Python. We'll use common libraries like pandas, statsmodels, and matplotlib for this.
1. Visual Checks (Quick & Intuitive)
You already started with the ACF plot, but let's expand on visualizations to confirm your initial observations:
a. Time Series Plot
First, plot the raw data to check if mean and variance stay consistent over time:
import pandas as pd import matplotlib.pyplot as plt # Assume your time series is stored as a pandas Series with datetime index speed_series = pd.read_csv("your_speed_data.csv", index_col=0, parse_dates=True) plt.figure(figsize=(12, 6)) plt.plot(speed_series) plt.title("Highway Speed Time Series") plt.xlabel("Time") plt.ylabel("Speed") plt.show()
Non-stationary data will show clear trends, seasonal spikes, or changing variance (e.g., speeds dropping during rush hours every day).
b. ACF & PACF Plots
You mentioned the ACF plot—non-stationary sequences typically have ACF values that decay very slowly (or not at all):
from statsmodels.graphics.tsaplots import plot_acf, plot_pacf fig, (ax1, ax2) = plt.subplots(2, 1, figsize=(12, 8)) plot_acf(speed_series, lags=48, ax=ax1) # 48 lags = 12 hours (15min intervals) plot_pacf(speed_series, lags=48, ax=ax2) plt.show()
For your seasonal data (daily cycles), you'll likely see strong peaks at lags 96 (1 full day) and multiples of that in the ACF plot.
2. Statistical Hypothesis Tests
Visual checks are great, but statistical tests give you a formal way to confirm stationarity. Here are the most widely used ones:
a. Augmented Dickey-Fuller (ADF) Test
Null Hypothesis (H₀): The time series is non-stationary (has a unit root).
Alternative Hypothesis (H₁): The time series is stationary.
If the p-value is less than 0.05, we reject H₀ and conclude the series is stationary.
from statsmodels.tsa.stattools import adfuller result = adfuller(speed_series) print(f"ADF Statistic: {result[0]:.4f}") print(f"p-value: {result[1]:.4f}") print("Critical Values:") for key, value in result[4].items(): print(f" {key}: {value:.4f}")
b. KPSS Test
Null Hypothesis (H₀): The time series is stationary (trend-stationary).
Alternative Hypothesis (H₁): The time series is non-stationary.
This complements the ADF test—if KPSS returns a p-value <0.05, we reject H₀ and conclude non-stationarity.
from statsmodels.tsa.stattools import kpss result = kpss(speed_series, regression='c') # 'c' for constant, 'ct' for trend+constant print(f"KPSS Statistic: {result[0]:.4f}") print(f"p-value: {result[1]:.4f}") print("Critical Values:") for key, value in result[3].items(): print(f" {key}: {value:.4f}")
c. Phillips-Perron (PP) Test
Similar to ADF but more robust to heteroscedasticity (changing variance) in the data:
from statsmodels.tsa.stattools import phillips_perron result = phillips_perron(speed_series) print(f"PP Statistic: {result[0]:.4f}") print(f"p-value: {result[1]:.4f}") print("Critical Values:") for key, value in result[4].items(): print(f" {key}: {value:.4f}")
3. Handling Seasonality-Specific Stationarity
Since your data has a clear daily seasonal component (96 observations per day), you should also test stationarity after accounting for seasonality:
a. Seasonal Decomposition
First, decompose the series into trend, seasonal, and residual components, then test the residuals for stationarity:
from statsmodels.tsa.seasonal import seasonal_decompose decomposition = seasonal_decompose(speed_series, model='additive', period=96) # period=96 for daily seasonality residuals = decomposition.resid.dropna() # Test residuals with ADF result = adfuller(residuals) print(f"ADF Test on Residuals - p-value: {result[1]:.4f}")
If the residuals are stationary, that means removing seasonality and trend leaves a stationary series.
b. Seasonal Difference
Another approach is to apply seasonal differencing (subtract the value from the same time one period ago) and test the differenced series:
seasonal_diff = speed_series.diff(periods=96).dropna() # Test the differenced series with ADF result = adfuller(seasonal_diff) print(f"ADF Test on Seasonally Differenced Data - p-value: {result[1]:.4f}")
For seasonal data, this often makes the series stationary, which is a key step before fitting your AR(p) model (or consider SARIMA instead, which is built for seasonal time series).
Final Tips
- Always combine visual checks with statistical tests—no single method is perfect.
- For your seasonal data, don't stop at testing the raw series: test residuals after decomposition or seasonally differenced data to get a clear picture.
- If tests confirm non-stationarity, apply differencing (regular and/or seasonal) before fitting your AR(p) model.
内容的提问来源于stack exchange,提问作者MehdiOua

