R语言forecast预测公路车流量结果远超出实际数据的问题求助
R语言forecast预测公路车流量结果远超出实际数据的问题求助
大家好,我现在在做公路车流量的预测工作,遇到了一个很棘手的问题:用同样的代码给部分路段做预测时结果正常,但针对某一路段,预测结果居然到了万亿级别,而实际最后一次观测的车流量(Leves)只有2,674,496,完全不符合预期。
我的建模流程简述
- 为了适配模型需求,我对数据做了差分处理(
diff(x)) - 使用
dLagM包的ardlDlm函数建模,阶数是老板之前用EViews估算好的;模型核心是解释PIB对Leves的影响,其他变量都是虚拟变量,所以移除了它们的滞后项 - 模型的拟合结果看起来合理,符合预期,但一到预测环节就出现了异常的万亿级结果
数据样本(前12行)
| Date | Leves | Pesados | Q1 | Q2 | PIB | JAN | FEV | MAR | ABR | MAI | JUN | JUL | AGO | SET | OUT | NOV | Crise | Greve | CES |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 01/04/1999 | 1106640.5 | 758932 | 1 | 0 | 60123.20 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 01/05/1999 | 1218219.5 | 827682 | 1 | 0 | 60957.24 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 01/06/1999 | 1109770.5 | 774122 | 1 | 0 | 62562.70 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 01/07/1999 | 1222242.5 | 757961 | 1 | 0 | 61539.71 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 01/08/1999 | 1078172.5 | 880780 | 1 | 0 | 61903.36 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| 01/09/1999 | 1120372.5 | 863346 | 1 | 0 | 60594.21 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 |
| 01/10/1999 | 1149853.5 | 869919 | 1 | 0 | 63358.14 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 |
| 01/11/1999 | 1094217.0 | 818265 | 1 | 0 | 65177.98 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 |
| 01/12/1999 | 1227005.0 | 821046 | 1 | 0 | 64274.77 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 01/01/2000 | 1160712.5 | 780890 | 1 | 0 | 59871.67 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 01/02/2000 | 1015884.0 | 794188 | 1 | 0 | 59273.23 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 01/03/2000 | 1152619.0 | 843165 | 1 | 0 | 59664.81 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
我的代码
library(openxlsx) library(egcm) library(aTSA) library(nortest) library(data.table) library(dynamac) library(AER) library(dynlm) library(readxl) library(stargazer) library(scales) library(urca) library(dLagM) library(dplyr) library(ggfortify) library(data.table) library(nardl) library(zoo) library(vars) library(neuralnet) library(DMwR2) library(TTR) library(quantmod) library(PerformanceAnalytics) library(ggplot2) library(writexl) library(tseries) library(ecm) library(knitr) library(plotly) library(tictoc) library(ARDL) #DADOS ---- Renovias <- read_excel("Ecopistas/ABCR/data.xlsx", sheet = "Renovias", range = "A1:t298") data <- Renovias[c(1:(nrow(Renovias)-48)),-1] d_Leves = diff(data$Leves) d_Pesados = diff(data$Pesados) d_PIB = diff(data$PIB) d_data = data[-1,] d_data$Leves = d_Leves d_data$Pesados = d_Pesados d_data$PIB = d_PIB #LEVES ------ newdata_leves <- Renovias[c((nrow(data)):nrow(Renovias)), c(4:19)] PIBdn = diff(newdata_leves$PIB) newdata_leves = newdata_leves[-1,] newdata_leves$PIB = PIBdn transposed_newdata_leves <- t(newdata_leves) model_leves = dLagM::ardlDlm(formula = Leves ~ PIB + JAN + FEV + MAR + ABR + MAI + JUN + JUL + AGO + SET + OUT + NOV + Crise + Greve + Q1 + Q2, data = d_data, p = 1, q = 2, remove = list(p = list(JAN=c(1:1),FEV=c(1:1),MAR=c(1:1), ABR=c(1:1), MAI=c(1:1), JUN=c(1:1),JUL=c(1:1), AGO=c(1:1), SET=c(1:1), OUT=c(1:1),NOV=c(1:1), Crise = c(1:1), Greve=c(1:1), Q1 = c(1:1), Q2 = c(1:1) ))) summary(model_leves) MAPE(model_leves) model_leves[["model"]][["coefficients"]][is.na(model_leves[["model"]][["coefficients"]])] <- 0 prediction_leves <- forecast(model = model_leves, x = transposed_newdata_leves, h = nrow(newdata_leves), interval = TRUE, level = 0.95, nSim = 1000)
求助思路
我目前完全找不到问题所在,有没有大佬能给点排查方向?比如:
- 是不是差分后的预测结果没有还原回原始尺度?毕竟建模用的是
diff(Leves),预测出来的是不是增量,需要累加回最后一个原始观测值? - 新数据的处理有没有问题?比如
newdata_leves的索引是否正确,差分后的PIB和建模数据的时序衔接是否连续? - 转置新数据传给
forecast函数是不是格式错了?ardlDlm的预测函数对输入x的格式有没有特定要求? - 虚拟变量在新数据中的取值是否正确?有没有可能某个变量被错误编码导致预测异常?
- 把模型的NA系数设为0的操作是否合理?会不会引入偏差影响预测?
备注:内容来源于stack exchange,提问作者Beatriz Roumanos
相关产品推荐
相关产品推荐

