基于R语言的百万级层级分组时间序列预测技术咨询
Great question! Given your dataset characteristics—1M monthly time series, each with 36 observations (3 years of data), needing 12-step forecasts, plus hierarchical (country → region) and grouped (product categories) structure—the gts package is absolutely a strong candidate to consider. Let's break down why, and walk through the technical implementation path you can follow:
Why gts is a good fit
gts(short for Hierarchical and Grouped Time Series) was built specifically for scenarios like yours where time series have cross-sectional hierarchies or groupings. Since each individual series only has 36 observations (relatively short for time series forecasting), leveraging aggregated information from higher levels (e.g., regional totals, product group totals) will almost certainly improve forecast accuracy compared to forecasting each series in isolation.- It handles both pure hierarchical structures (e.g., Region → Country → Product) and grouped structures (e.g., product categories cutting across countries), as well as combined "cross-sectional" hierarchies where both dimensions overlap.
Technical Implementation Path
Let’s break this down into actionable steps:
1. Data Preparation
First, you’ll need to structure your data to work with gts:
- Start by formatting your data into a matrix or
mts(multivariate time series) object where each column represents a bottom-level series (e.g., each country-product combination), and each row represents a month. - Define your hierarchy/grouping structure using a list. For example, if you have regions containing countries, and products grouped into categories, you’d create a structure that maps bottom-level series to their parent groups. The
gts()function will use this to establish the relationships between series.# Example hierarchy definition (adjust to your actual data structure) hierarchy <- list( region = c("RegionA", "RegionA", "RegionB", "RegionB"), # Maps each bottom series to its region product_group = c("GroupX", "GroupY", "GroupX", "GroupY") # Maps to its product group ) # Convert your time series matrix to a gts object my_gts <- gts(data = my_time_series_matrix, groups = hierarchy)
2. Generate Base Forecasts
gts doesn’t handle forecasting directly—it relies on base forecasts for each bottom-level series, then reconciles them to fit the hierarchy. For 1M series, you’ll want an efficient, automated method:
- Use
auto.arima()from theforecastpackage (the standard for automated ARIMA) orets()for exponential smoothing. Both work well, butets()is faster for large datasets. - For even faster performance, consider
thetaf()(Theta method), which is designed for short time series and runs quickly. - Parallelize this step: With 1M series, sequential processing will be far too slow. Use packages like
foreachanddoParallelto distribute the forecasting across multiple cores:library(foreach) library(doParallel) cl <- makeCluster(detectCores() - 1) # Leave one core free for system tasks registerDoParallel(cl) # Generate base forecasts for all bottom-level series base_forecasts <- foreach(i = 1:ncol(my_time_series_matrix), .packages = "forecast") %dopar% { auto.arima(my_time_series_matrix[, i]) %>% forecast(h = 12) } stopCluster(cl)
3. Reconcile Forecasts
Once you have base forecasts for all bottom-level series, use gts to reconcile them so that the forecasts align with your hierarchy (e.g., regional forecasts equal the sum of their countries' forecasts):
- The
combinef()function ingtshandles this. You can choose from several reconciliation methods:ols: Ordinary Least Squares (simple, but assumes equal forecast error variance)wls: Weighted Least Squares (weights by inverse variance of base forecasts)mint: Minimum Trace (often the best-performing, as it accounts for forecast error covariance)
The# Convert base forecasts to a matrix of forecast means fc_matrix <- sapply(base_forecasts, function(x) x$mean) # Reconcile forecasts across all hierarchy levels reconciled_fc <- combinef(fc_matrix, my_gts$groups, method = "mint", keep = "all")keep = "all"argument gives you forecasts for every level of the hierarchy (bottom-level, country, region, product group, total), which is useful if you need forecasts at multiple levels.
4. Alternatives to Consider
If gts feels constrained, or you want to work within the tidyverse ecosystem:
fabletools(tidyverts): This modern package supports hierarchical forecast reconciliation with a more intuitive, tidy workflow. You can define hierarchies usinghierarchy_tree()and reconcile forecasts withreconcile(). It works seamlessly withfable's forecast methods (likeARIMA()orETS()).htspackage: A predecessor togts, focused specifically on hierarchical (not grouped) time series. Still solid, butgtsis more flexible for your grouped structure.- For extremely large datasets, you might also explore lightweight methods like
forecastHybrid(combining multiple base forecasts) or even machine learning approaches (e.g., LSTMs), but note that ML methods require more work to integrate hierarchical reconciliation compared to statistical tools likegts.
Key Considerations
- Computational Resources: 1M series will require significant memory and processing power. Parallelization is non-negotiable here. You might also consider sampling a subset of your data first to test your workflow before scaling to the full dataset.
- Model Selection: Test a few base forecast methods on a sample to see which gives the best performance for your data. Short series (36 observations) often do well with exponential smoothing or Theta methods, which are faster than ARIMA.
内容的提问来源于stack exchange,提问作者pk_22

