You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言时间序列‘Subscript out of bounds’错误的解决方法

问题背景

使用美国劳工统计局的就业时间序列数据(前10行如下):

> Time_Series
# A tsibble: 115 x 8 [1M]
    year month  value Year  Month Date1        Date  Value
   <int> <int>  <dbl> <chr> <chr> <chr>       <mth>  <dbl>
 1  2016     1 143210 2016  Jan   2016 Jan 2016 Jan 143210
 2  2016     2 143407 2016  Feb   2016 Feb 2016 Feb 143407
 3  2016     3 143662 2016  Mar   2016 Mar 2016 Mar 143662
 4  2016     4 143855 2016  Apr   2016 Apr 2016 Apr 143855
 5  2016     5 143900 2016  May   2016 May 2016 May 143900
 6  2016     6 144146 2016  Jun   2016 Jun 2016 Jun 144146
 7  2016     7 144520 2016  Jul   2016 Jul 2016 Jul 144520
 8  2016     8 144662 2016  Aug   2016 Aug 2016 Aug 144662
 9  2016     9 144967 2016  Sep   2016 Sep 2016 Sep 144967
10  2016    10 145066 2016  Oct   2016 Oct 2016 Oct 145066
# ℹ 105 more rows

将数据拆分为训练集与测试集:

Time_Series_train <- structure(list(Date = structure(c(16832, 16861, 16892, 16922, 
16953, 16983, 17014, 17045, 17075, 17106, 17136, 17167, 17198, 
17226, 17257, 17287, 17318, 17348, 17379, 17410, 17440, 17471, 
17501, 17532, 17563, 17591, 17622, 17652, 17683, 17713, 17744, 
17775, 17805, 17836, 17866, 17897, 17928, 17956, 17987, 18017, 
18048, 18078, 18109, 18140, 18170, 18201, 18231, 18262, 18293, 
18322, 18353, 18383, 18414, 18444, 18475, 18506, 18536, 18567, 
18597, 18628, 18659, 18687, 18718, 18748, 18779, 18809, 18840, 
18871), class = c("yearmonth", "vctrs_vctr")), Value = c(197, 
255, 193, 45, 246, 374, 142, 305, 99, 117, 225, 220, 220, 121, 
205, 206, 203, 189, 147, 88, 143, 223, 150, 137, 394, 226, 142, 
318, 219, 61, 259, 79, 168, 91, 192, 250, 6, 230, 298, 28, 218, 
97, 235, 194, 95, 208, 127, 236, 261, -1397, -20471, 2616, 4631, 
1584, 1564, 951, 691, 270, -183, 365, 509, 824, 365, 421, 796, 
931, 487, 466)), class = c("tbl_ts", "tbl_df", "tbl", "data.frame"
), row.names = c(NA, -68L), key = structure(list(.rows = structure(list(
    1:68), ptype = integer(0), class = c("vctrs_list_of", "vctrs_vctr", 
"list"))), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, 
-1L)), index = structure("Date", ordered = TRUE), index2 = "Date", interval = structure(list(
    year = 0, quarter = 0, month = 1, week = 0, day = 0, hour = 0, 
    minute = 0, second = 0, millisecond = 0, microsecond = 0, 
    nanosecond = 0, unit = 0), .regular = TRUE, class = c("interval", 
"vctrs_rcrd", "vctrs_vctr")))
Time_Series_test <- structure(list(Date = structure(c(18901, 18932, 18962, 18993, 
19024, 19052, 19083, 19113, 19144, 19174, 19205, 19236, 19266, 
19297, 19327, 19358, 19389, 19417, 19448, 19478, 19509, 19539, 
19570, 19601, 19631, 19662, 19692, 19723, 19754, 19783, 19814, 
19844, 19875, 19905, 19936, 19967, 19997, 20028, 20058, 20089, 
20120, 20148, 20179, 20209, 20240), class = c("yearmonth", "vctrs_vctr"
)), Value = c(857, 637, 575, 225, 869, 471, 305, 241, 461, 696, 
237, 227, 400, 297, 126, 444, 306, 85, 216, 227, 257, 148, 157, 
158, 186, 141, 269, 119, 222, 246, 118, 193, 87, 88, 71, 240, 
44, 261, 323, 111, 102, 120, 158, 19, 14)), class = c("tbl_ts", 
"tbl_df", "tbl", "data.frame"), row.names = c(NA, -45L), key = structure(list(
    .rows = structure(list(1:45), ptype = integer(0), class = c("vctrs_list_of", 
    "vctrs_vctr", "list"))), class = c("tbl_df", "tbl", "data.frame"
), row.names = c(NA, -1L)), index = structure("Date", ordered = TRUE), index2 = "Date", interval = structure(list(
    year = 0, quarter = 0, month = 1, week = 0, day = 0, hour = 0, 
    minute = 0, second = 0, millisecond = 0, microsecond = 0, 
    nanosecond = 0, unit = 0), .regular = TRUE, class = c("interval", 
"vctrs_rcrd", "vctrs_vctr")))

模型构建

基于fpp3包构建4个TSLM线性时间序列模型:

library(fpp3)
fit <- Time_Series_train %>%
  fabletools::model(
    Linear1 = fable::TSLM(Value ~ season() + trend()),
    Linear2 = fable::TSLM(Value),
    Linear3 = fable::TSLM(Value ~ season()),
    Linear4 = fable::TSLM(Value ~ trend())
  )

报错情况

执行以下精度评估代码时:

Value_forecast_accuracy <- fit %>%
  fabletools::forecast(h = 1) %>%
  fabletools::accuracy(Time_Series_test) %>%
  select(.model:MAPE) %>%
  arrange(RMSE)

虽得到精度结果:

# A tibble: 4 × 7
  .model  .type    ME  RMSE   MAE   MPE  MAPE
  <chr>   <chr> <dbl> <dbl> <dbl> <dbl> <dbl>
1 Linear3 Test   618.  618.  618.  72.1  72.1
2 Linear1 Test   769.  769.  769.  89.7  89.7
3 Linear2 Test   791.  791.  791.  92.3  92.3
4 Linear4 Test   877.  877.  877. 102.  102. 

但出现4次“subscript out of bounds”(下标越界)警告,如何解决该错误?


解决方法

错误原因

警告的核心原因是:forecast(h=1)仅生成了1步预测(即仅1个时间点的预测值),但测试集包含45个观测。accuracy()函数尝试将预测值与测试集的所有45个实际值匹配时,会试图访问超出预测结果范围的索引,从而触发下标越界警告。

修正代码

生成与测试集长度一致的预测,确保预测的时间范围和测试集完全对齐:

Value_forecast_accuracy <- fit %>%
  fabletools::forecast(h = nrow(Time_Series_test)) %>%  # 生成与测试集行数匹配的预测
  fabletools::accuracy(Time_Series_test) %>%
  select(.model:MAPE) %>%
  arrange(RMSE)

说明

  • 使用nrow(Time_Series_test)替代固定的h=1,可以自动匹配测试集的观测数量,生成对应步数的预测。
  • 修正后,预测结果的时间点会和测试集完全衔接,accuracy()能正确匹配每个时间点的预测值与实际值,不会再出现下标越界警告。
  • 同时,得到的精度指标是基于整个测试集的完整评估,比仅1步预测的结果更具参考价值。

内容的提问来源于stack exchange,提问作者Russ Conte

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 10:22:01