咨询VAR模型小时数据的滞后阶数选择(基于5年小时级数据)
Great question—hourly time series bring unique cyclical patterns that don’t apply to quarterly or monthly data, so standard rules of thumb don’t directly translate. Let’s walk through the key steps to pick a reasonable lag order for your VAR model:
1. Start with Intrinsic Cyclical Lags
Hourly data has strong natural cycles you can’t ignore:
- Daily cycle: 24 hours (a full day of observations). At minimum, you’ll want to include lags up to 24 to capture day-to-day repeating patterns (e.g., hourly activity fluctuations that mirror each day).
- Weekly cycle: 168 hours (7*24). If your data shows weekly trends (like lower activity on weekends vs. weekdays), lags around 168 will be critical to model that behavior.
- Longer cycles (monthly/quarterly) might exist too, but they’re often less impactful at the hourly level unless your data has clear monthly seasonality.
2. Use Statistical Information Criteria
Standard criteria help balance model fit and parsimony—run these to narrow down your candidate lag range:
- AIC (Akaike Information Criterion): Tends to favor more lags, useful if you want to capture subtle patterns.
- BIC (Bayesian Information Criterion): Penalizes complex models more heavily, leaning toward fewer lags to avoid overfitting.
- HQIC (Hannan-Quinn Information Criterion): A middle ground between AIC and BIC.
In practice, you can calculate these using common tools:
- R:
vars::VARselect(your_data, lag.max = 200)(set a max lag high enough to cover weekly cycles) - Python:
statsmodels.tsa.vector_ar.var_model.VAR.select_order(your_data, maxlags=200)
Don’t just pick the top-ranked criterion—look for consensus across multiple criteria, or note where they diverge to dig deeper.
3. Validate with Residual White Noise Test
No matter which lag order you pick, confirm the model’s residuals are white noise (no remaining autocorrelation):
- Run a Ljung-Box test on the residuals of each variable in the VAR. If the test fails to reject the null hypothesis (p-value > 0.05), your lags are sufficient to capture the data’s autocorrelation.
- If residuals still show autocorrelation, you’ll need to increase the lag order or add cyclical lags you might have missed (e.g., 48 hours for two-day patterns).
4. Account for Degrees of Freedom
You have ~43,800 hourly observations (5 years * 365 days * 24 hours), which gives you more flexibility than monthly/quarterly data—but don’t overdo it:
- A VAR with k variables and p lags has k²p parameters. For example, 3 variables with 168 lags mean 3²168 = 1,512 parameters—manageable with your sample size, but this can get unwieldy if you have more variables.
- If high lags cause overfitting, consider constrained VARs: only include lags that align with meaningful cycles (e.g., 24, 48, 168) instead of every lag up to p. Sparse VAR methods can also help reduce unnecessary parameters.
5. Test Robustness
Try 2-3 reasonable lag orders (e.g., 24, 48, 168, or the top picks from information criteria) and compare key outputs:
- Do impulse response functions look stable and consistent across models?
- Does variance decomposition show similar patterns of variable influence?
- If results don’t change drastically, your lag order choice is robust. If they do, dig into why—maybe one order is missing a critical cyclical pattern.
内容的提问来源于stack exchange,提问作者junmouse

