You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于scikit-hts的分组时间序列销量预测报错及多序列方案咨询

多时间序列销量预测相关问题

我正在对Kaggle的Store item demand forecasting challenge数据集做多时间序列销量预测,该数据集包含10家门店、50个商品对应的500条堆叠存储的长格式时间序列,每个门店-商品组合有5年日度记录,同时存在周度和年度季节性。

总数据量为:365.2天 * 5年 * 10门店 *50商品 = 913000条记录。

根据我此前阅读的Hierarchical and Grouped time series相关知识,整个数据集可构建为分组时间序列(Grouped Time Series)而非严格的层级时间序列,因为支持按门店、商品两个维度灵活聚合。

我希望使用scikit-hts库的AutoArimaModel函数(pmdarima的AutoArima封装函数)预测全部500条时间序列(store1_item1、store1_item2……store10_item50)2015年1月1日至12月31日的全年数据。

为处理两层季节性,我添加傅里叶项作为外生特征处理年度季节性,周度季节性由auto_arima自行处理。

目前我首先在预测阶段遇到报错:

ValueError: Provided exogenous values are not of the appropriate shape. Required (365, 4), got (365, 8).

我判断是外生变量字典配置有误,但我是首次使用scikit-hts,按照官方文档操作仍未解决该问题。

编辑补充

我后续发现Github上有同类问题报告,按照给出的方案本地修复后无报错,但出现新的问题:部分预测结果为负,正值预测结果也存在数值失真的问题,仅少数门店-商品组合的预测结果符合预期,相关绘图示例如下:

df.loc['2014','store_1_item_1'].plot()
predictions.loc['2015','store_1_item_1'].plot()

store1_item1预测结果图

df.loc['2014','store_1_item_2'].plot()
predictions.loc['2015','store_1_item_2'].plot()

store1_item2预测结果图

df.loc['2014','store_2_item_1'].plot()
predictions.loc['2015','store_2_item_1'].plot()

store2_item1预测结果图

df.loc['2014','store_2_item_2'].plot()
predictions.loc['2015','store_2_item_2'].plot()

store2_item2预测结果图

完整代码

# imports
import pandas as pd
from pmdarima.preprocessing import FourierFeaturizer
import hts
from hts.hierarchy import HierarchyTree
from hts.model import AutoArimaModel
from hts import HTSRegressor


# read data from the csv file
data = pd.read_csv('train.csv', index_col='date', parse_dates=True)

# Train/Test split with reduced size
train_data = data.query('store == [1,2] and item == [1, 2]').loc['2013':'2014']
test_data = data.query('store == [1,2] and item == [1, 2]').loc['2015']


# Create the stores time series
# For each timestamp group by store and apply sum
stores_ts = train_data.drop(columns=['item']).groupby(['date','store']).sum()
stores_ts = stores_ts.unstack('store')
stores_ts.columns = stores_ts.columns.droplevel(0)
stores_ts.columns = ['store_' + str(i) for i in stores_ts.columns]

# Create the items time series
# For each timestamp group by item and apply sum
items_ts = train_data.drop(columns=['store']).groupby(['date','item']).sum()
items_ts = items_ts.unstack('item')
items_ts.columns = items_ts.columns.droplevel(0)
items_ts.columns = ['item_' + str(i) for i in items_ts.columns]


# Create the stores_items time series
# For each timestamp group by store AND by item and apply sum
store_item_ts = train_data.pivot_table(index= 'date', columns=['store', 'item'], aggfunc='sum')
store_item_ts.columns = store_item_ts.columns.droplevel(0)

# Rename the columns as store_i_item_j
col_names = []
for i in store_item_ts.columns:
    col_name = 'store_' + str(i[0]) + '_item_' + str(i[1])
    col_names.append(col_name)
    
store_item_ts.columns = store_item_ts.columns.droplevel(0)
store_item_ts.columns = col_names

# Create a new dataframe and add the root level of the hierarchy as the sum of all stores (or all items)
df = pd.DataFrame()
df['total'] = stores_ts.sum(1) 

# Concatenate all created dataframes into one df
# df is the dataframe that will be used for model training
df = pd.concat([df, stores_ts, items_ts, store_item_ts], 1)


# Build fourier terms for train and test sets
four_terms = FourierFeaturizer(365.2, 1)

# Build the exogenous features dataframe for training data
exog_train_df = pd.DataFrame()

for i in range(1, 3):
    for j in range(1, 3):
        _, exog = four_terms.fit_transform(train_data.query(f'store == {i} and item == {j}').sales)
        exog.columns= [f'store_{i}_item_{j}_'+ x for x in exog.columns]
        exog_train_df = pd.concat([exog_train_df, exog], axis=1)
exog_train_df['date'] = df.index
exog_train_df.set_index('date', inplace=True)

# add the exogenous features dataframe to df before training
df = pd.concat([df, exog_train_df], axis= 1)


# Build the exogenous features dataframe for test set
# It will be used only when using model.predict()
exog_test_df = pd.DataFrame()

for i in range(1, 3):
    for j in range(1, 3):
        _, exog_test = four_terms.fit_transform(test_data.query(f'store == {i} and item == {j}').sales)
        exog_test.columns= [f'store_{i}_item_{j}_'+ x for x in exog_test.columns]
        exog_test_df = pd.concat([exog_test_df, exog_test], axis=1)


# Build the hierarchy of the Grouped Time Series
stores = [i for i in stores_ts.columns]
items = [i for i in items_ts.columns]
store_items = col_names

# Exogenous features mapping
exog_store_items = {e: [v for v in exog_train_df.columns if v.startswith(e)] for e in store_items}  
exog_stores = {e:[v for v in exog_train_df.columns if v.startswith(e)] for e in stores}
exog_items = {e:[v for v in exog_train_df.columns if v.find(e) != -1] for e in items}
exog_total = {'total':[v for v in exog_train_df.columns if v.find('FOURIER') != -1]}

# Merge all dictionaries
exog_to_merge = [exog_store_items, exog_stores, exog_items, exog_total]
exogenous = {k:v for x in exog_to_merge for k,v in x.items()}

# Build hierarchy
total = {'total': stores + items}
store_h = {k: [v for v in store_items if v.startswith(k)] for k in stores}
hierarchy = {**total, **store_h}

# Hierarchy tree automatically created by hts
ht = HierarchyTree.from_nodes(nodes=hierarchy, df=df, exogenous=exogenous)

# Instanciate the auto arima model using HTSRegressor
autoarima = HTSRegressor(model='auto_arima', D=1, m=7, seasonal=True, revision_method='OLS', n_jobs=12)

# Fit the model to the training df that includes time series and exog_train_df
# Set exogenous param to the previously built dictionary
model = autoarima.fit(df, hierarchy, exogenous=exogenous)

# Make predictions
# Set the exogenous_df param 
predictions = model.predict(exogenous_df=exog_test_df, steps_ahead=365)

我还尝试过另外两种方案,在单条时间序列(如store1_item1)上验证有效:

  • 遍历500条时间序列,对每条独立使用TBATS建模预测
  • 遍历500条时间序列,对每条独立使用带傅里叶外生特征的auto_arima(SARIMAX)建模预测

请问上述两种方案是否合理?有没有更适合多时间序列场景的ARIMA扩展方案?
另外我想尝试用LSTM完成该任务,我是数据科学和深度学习领域新手,请问数据预处理应该采用原始长格式对门店、商品字段做独热编码,还是使用本次处理后的宽表df?


问题解答

现有方案合理性判断

你提到的两种单序列独立建模的方案完全合理,适配当前场景的优势很明确:

  • TBATS原生支持多周期季节性,不需要额外手动构造傅里叶外生变量,代码实现更简洁,对销量的周、年双季节性适配性好,缺点是500条序列并行训练耗时略高于SARIMAX。
  • 带傅里叶项的auto_arima方案可以灵活控制季节性拟合的复杂度,你已经在单序列上验证有效,批量跑的时候只要做好异常捕获、负值截断即可落地。

针对多时间序列场景的ARIMA类扩展方案,可以选择全局ARIMA(GlobalARIMA)思路:所有序列共享相同的ARIMA阶数参数,仅拟合各自的截距和外生变量系数,相比单序列独立建模训练速度快2~3倍,预测精度损失极小,非常适合这类同分布的门店-商品销量预测场景,可以用pmdarima的批量建模接口结合自定义全局参数搜索逻辑实现。

scikit-hts预测失真的原因

你遇到的负值、预测失真问题核心是层级调和逻辑和外生变量配置冲突:

  1. 你构造的外生变量映射有冗余,上层节点(比如store_1、total节点)的外生变量是多个底层序列傅里叶项的拼接,维度和模型训练阶段期望的输入不匹配,即使修复了Github issue提到的代码bug,模型拟合的参数本身就是错误的。
  2. OLS调和方法会强制所有层级的预测结果满足聚合约束,当底层序列预测本身有偏差时,调和过程会进一步放大误差,甚至出现负值。如果坚持用scikit-hts,可以先把调和方法换成'BU'(BottomUp,底层预测后直接向上聚合),先验证底层序列预测的准确性,再考虑更复杂的调和策略。

LSTM数据预处理建议

作为新手优先选择原始长格式+独热编码的方案,优势如下:

  • 长格式输入的结构更符合LSTM序列化输入的逻辑,每个样本对应[时间步, 特征数]的维度,特征可以包含独热编码后的门店ID、商品ID、历史销量、傅里叶季节性特征、日期衍生特征(是否节假日、周几、月份等)。
  • 宽表的列数会随着门店、商品数量增加暴增,输入维度太高很容易导致LSTM过拟合,训练难度大,对新手不友好。
  • 长格式可以很方便地做滑动窗口构造训练样本,也支持未来扩展全局共享的特征,泛化性更好。

内容的提问来源于stack exchange,提问作者Downforu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 09:00:03