You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言的百万级层级分组时间序列预测技术咨询

Great question! Given your dataset characteristics—1M monthly time series, each with 36 observations (3 years of data), needing 12-step forecasts, plus hierarchical (country → region) and grouped (product categories) structure—the gts package is absolutely a strong candidate to consider. Let's break down why, and walk through the technical implementation path you can follow:

Why gts is a good fit

  • gts (short for Hierarchical and Grouped Time Series) was built specifically for scenarios like yours where time series have cross-sectional hierarchies or groupings. Since each individual series only has 36 observations (relatively short for time series forecasting), leveraging aggregated information from higher levels (e.g., regional totals, product group totals) will almost certainly improve forecast accuracy compared to forecasting each series in isolation.
  • It handles both pure hierarchical structures (e.g., Region → Country → Product) and grouped structures (e.g., product categories cutting across countries), as well as combined "cross-sectional" hierarchies where both dimensions overlap.

Technical Implementation Path

Let’s break this down into actionable steps:

1. Data Preparation

First, you’ll need to structure your data to work with gts:

  • Start by formatting your data into a matrix or mts (multivariate time series) object where each column represents a bottom-level series (e.g., each country-product combination), and each row represents a month.
  • Define your hierarchy/grouping structure using a list. For example, if you have regions containing countries, and products grouped into categories, you’d create a structure that maps bottom-level series to their parent groups. The gts() function will use this to establish the relationships between series.
    # Example hierarchy definition (adjust to your actual data structure)
    hierarchy <- list(
      region = c("RegionA", "RegionA", "RegionB", "RegionB"), # Maps each bottom series to its region
      product_group = c("GroupX", "GroupY", "GroupX", "GroupY") # Maps to its product group
    )
    # Convert your time series matrix to a gts object
    my_gts <- gts(data = my_time_series_matrix, groups = hierarchy)
    

2. Generate Base Forecasts

gts doesn’t handle forecasting directly—it relies on base forecasts for each bottom-level series, then reconciles them to fit the hierarchy. For 1M series, you’ll want an efficient, automated method:

  • Use auto.arima() from the forecast package (the standard for automated ARIMA) or ets() for exponential smoothing. Both work well, but ets() is faster for large datasets.
  • For even faster performance, consider thetaf() (Theta method), which is designed for short time series and runs quickly.
  • Parallelize this step: With 1M series, sequential processing will be far too slow. Use packages like foreach and doParallel to distribute the forecasting across multiple cores:
    library(foreach)
    library(doParallel)
    cl <- makeCluster(detectCores() - 1) # Leave one core free for system tasks
    registerDoParallel(cl)
    
    # Generate base forecasts for all bottom-level series
    base_forecasts <- foreach(i = 1:ncol(my_time_series_matrix), .packages = "forecast") %dopar% {
      auto.arima(my_time_series_matrix[, i]) %>% forecast(h = 12)
    }
    stopCluster(cl)
    

3. Reconcile Forecasts

Once you have base forecasts for all bottom-level series, use gts to reconcile them so that the forecasts align with your hierarchy (e.g., regional forecasts equal the sum of their countries' forecasts):

  • The combinef() function in gts handles this. You can choose from several reconciliation methods:
    • ols: Ordinary Least Squares (simple, but assumes equal forecast error variance)
    • wls: Weighted Least Squares (weights by inverse variance of base forecasts)
    • mint: Minimum Trace (often the best-performing, as it accounts for forecast error covariance)
    # Convert base forecasts to a matrix of forecast means
    fc_matrix <- sapply(base_forecasts, function(x) x$mean)
    
    # Reconcile forecasts across all hierarchy levels
    reconciled_fc <- combinef(fc_matrix, my_gts$groups, method = "mint", keep = "all")
    
    The keep = "all" argument gives you forecasts for every level of the hierarchy (bottom-level, country, region, product group, total), which is useful if you need forecasts at multiple levels.

4. Alternatives to Consider

If gts feels constrained, or you want to work within the tidyverse ecosystem:

  • fabletools (tidyverts): This modern package supports hierarchical forecast reconciliation with a more intuitive, tidy workflow. You can define hierarchies using hierarchy_tree() and reconcile forecasts with reconcile(). It works seamlessly with fable's forecast methods (like ARIMA() or ETS()).
  • hts package: A predecessor to gts, focused specifically on hierarchical (not grouped) time series. Still solid, but gts is more flexible for your grouped structure.
  • For extremely large datasets, you might also explore lightweight methods like forecastHybrid (combining multiple base forecasts) or even machine learning approaches (e.g., LSTMs), but note that ML methods require more work to integrate hierarchical reconciliation compared to statistical tools like gts.

Key Considerations

  • Computational Resources: 1M series will require significant memory and processing power. Parallelization is non-negotiable here. You might also consider sampling a subset of your data first to test your workflow before scaling to the full dataset.
  • Model Selection: Test a few base forecast methods on a sample to see which gives the best performance for your data. Short series (36 observations) often do well with exponential smoothing or Theta methods, which are faster than ARIMA.

内容的提问来源于stack exchange,提问作者pk_22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:36:32