You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R的通用时间序列回测与交叉验证技术实现问询

Rolling Time Series Evaluation Workflow in R

Got it, let's walk through how to implement this rolling time series evaluation approach in R—this is a rock-solid method for validating time series models without introducing look-ahead bias, which is crucial for real-world forecasting.

First, let's restate the core workflow you outlined to make sure we're on the same page:

评估周期规模与间隔为n,评估起始点为b,时间滞后为l;首先使用数据点1:b-l训练模型,在b:b+n区间进行评估;随后使用1:b+n-l训练模型,在b+n:b+2n区间评估,以此类推完成k个周期的评估。

This is often called rolling origin evaluation, and here's how to build it step by step:

Step 1: Set Up Parameters & Sample Data

First, let's define our key parameters and create a sample time series to test with (replace this with your actual dataset):

# Load handy libraries for time series and metrics
library(forecast)
library(Metrics)

# Generate sample monthly time series data (100 points)
set.seed(123)
ts_data <- ts(rnorm(100), frequency = 12)

# Define your evaluation parameters
n <- 6       # Size/interval of each evaluation period
b <- 24      # Starting point of the first evaluation period
l <- 3       # Time lag (exclude recent data before training cutoff)
k <- 5       # Total number of evaluation cycles

Step 2: Build the Rolling Evaluation Loop

We'll loop through each cycle, train the model on the correct historical data, forecast the evaluation window, and track metrics:

# Initialize a dataframe to store all evaluation results
eval_results <- data.frame(
  cycle = 1:k,
  train_data_end = integer(k),
  eval_period_start = integer(k),
  eval_period_end = integer(k),
  mae = numeric(k),
  rmse = numeric(k)
)

# Loop through each evaluation cycle
for (i in 0:(k-1)) {
  # Calculate indices for training and evaluation
  train_cutoff <- b + i*n - l
  eval_start <- b + i*n
  eval_end <- b + (i+1)*n - 1  # Adjust for 1-based indexing in R
  
  # Slice the data into training and evaluation subsets
  train_subset <- window(ts_data, end = train_cutoff)
  eval_subset <- window(ts_data, start = eval_start, end = eval_end)
  
  # Train your model (swap this with your model of choice: ARIMA, XGBoost, etc.)
  model <- auto.arima(train_subset)
  
  # Forecast the evaluation period
  forecast_vals <- forecast(model, h = n)$mean
  
  # Calculate and store metrics
  eval_results$train_data_end[i+1] <- train_cutoff
  eval_results$eval_period_start[i+1] <- eval_start
  eval_results$eval_period_end[i+1] <- eval_end
  eval_results$mae[i+1] <- mae(eval_subset, forecast_vals)
  eval_results$rmse[i+1] <- rmse(eval_subset, forecast_vals)
}

# View your results
print(eval_results)

Key Tips for Customization

  • Time Lag Purpose: The l parameter is perfect if you need to simulate a delay in data availability (e.g., you can't use the most recent 3 months of data to forecast the next period) or if your model relies on lagged features that need buffer space.
  • Model Flexibility: Replace auto.arima() with any model you're working with—use lm() for linear models with lagged predictors, xgboost() for gradient boosting, or even custom models.
  • Metric Adjustments: Swap or add metrics like MAPE, MASE, or R² based on what matters most for your use case.
  • Error Checking: Double-check that your dataset is long enough to cover b + k*n—if not, reduce k or adjust your parameters to avoid indexing errors.

内容的提问来源于stack exchange,提问作者gsmafra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:27:29