基于scikit-hts的分组时间序列销量预测报错及多序列方案咨询
我正在对Kaggle的Store item demand forecasting challenge数据集做多时间序列销量预测,该数据集包含10家门店、50个商品对应的500条堆叠存储的长格式时间序列,每个门店-商品组合有5年日度记录,同时存在周度和年度季节性。
总数据量为:365.2天 * 5年 * 10门店 *50商品 = 913000条记录。
根据我此前阅读的Hierarchical and Grouped time series相关知识,整个数据集可构建为分组时间序列(Grouped Time Series)而非严格的层级时间序列,因为支持按门店、商品两个维度灵活聚合。
我希望使用scikit-hts库的AutoArimaModel函数(pmdarima的AutoArima封装函数)预测全部500条时间序列(store1_item1、store1_item2……store10_item50)2015年1月1日至12月31日的全年数据。
为处理两层季节性,我添加傅里叶项作为外生特征处理年度季节性,周度季节性由auto_arima自行处理。
目前我首先在预测阶段遇到报错:
ValueError: Provided exogenous values are not of the appropriate shape. Required (365, 4), got (365, 8).
我判断是外生变量字典配置有误,但我是首次使用scikit-hts,按照官方文档操作仍未解决该问题。
编辑补充
我后续发现Github上有同类问题报告,按照给出的方案本地修复后无报错,但出现新的问题:部分预测结果为负,正值预测结果也存在数值失真的问题,仅少数门店-商品组合的预测结果符合预期,相关绘图示例如下:
df.loc['2014','store_1_item_1'].plot() predictions.loc['2015','store_1_item_1'].plot()

df.loc['2014','store_1_item_2'].plot() predictions.loc['2015','store_1_item_2'].plot()

df.loc['2014','store_2_item_1'].plot() predictions.loc['2015','store_2_item_1'].plot()

df.loc['2014','store_2_item_2'].plot() predictions.loc['2015','store_2_item_2'].plot()

完整代码
# imports import pandas as pd from pmdarima.preprocessing import FourierFeaturizer import hts from hts.hierarchy import HierarchyTree from hts.model import AutoArimaModel from hts import HTSRegressor # read data from the csv file data = pd.read_csv('train.csv', index_col='date', parse_dates=True) # Train/Test split with reduced size train_data = data.query('store == [1,2] and item == [1, 2]').loc['2013':'2014'] test_data = data.query('store == [1,2] and item == [1, 2]').loc['2015'] # Create the stores time series # For each timestamp group by store and apply sum stores_ts = train_data.drop(columns=['item']).groupby(['date','store']).sum() stores_ts = stores_ts.unstack('store') stores_ts.columns = stores_ts.columns.droplevel(0) stores_ts.columns = ['store_' + str(i) for i in stores_ts.columns] # Create the items time series # For each timestamp group by item and apply sum items_ts = train_data.drop(columns=['store']).groupby(['date','item']).sum() items_ts = items_ts.unstack('item') items_ts.columns = items_ts.columns.droplevel(0) items_ts.columns = ['item_' + str(i) for i in items_ts.columns] # Create the stores_items time series # For each timestamp group by store AND by item and apply sum store_item_ts = train_data.pivot_table(index= 'date', columns=['store', 'item'], aggfunc='sum') store_item_ts.columns = store_item_ts.columns.droplevel(0) # Rename the columns as store_i_item_j col_names = [] for i in store_item_ts.columns: col_name = 'store_' + str(i[0]) + '_item_' + str(i[1]) col_names.append(col_name) store_item_ts.columns = store_item_ts.columns.droplevel(0) store_item_ts.columns = col_names # Create a new dataframe and add the root level of the hierarchy as the sum of all stores (or all items) df = pd.DataFrame() df['total'] = stores_ts.sum(1) # Concatenate all created dataframes into one df # df is the dataframe that will be used for model training df = pd.concat([df, stores_ts, items_ts, store_item_ts], 1) # Build fourier terms for train and test sets four_terms = FourierFeaturizer(365.2, 1) # Build the exogenous features dataframe for training data exog_train_df = pd.DataFrame() for i in range(1, 3): for j in range(1, 3): _, exog = four_terms.fit_transform(train_data.query(f'store == {i} and item == {j}').sales) exog.columns= [f'store_{i}_item_{j}_'+ x for x in exog.columns] exog_train_df = pd.concat([exog_train_df, exog], axis=1) exog_train_df['date'] = df.index exog_train_df.set_index('date', inplace=True) # add the exogenous features dataframe to df before training df = pd.concat([df, exog_train_df], axis= 1) # Build the exogenous features dataframe for test set # It will be used only when using model.predict() exog_test_df = pd.DataFrame() for i in range(1, 3): for j in range(1, 3): _, exog_test = four_terms.fit_transform(test_data.query(f'store == {i} and item == {j}').sales) exog_test.columns= [f'store_{i}_item_{j}_'+ x for x in exog_test.columns] exog_test_df = pd.concat([exog_test_df, exog_test], axis=1) # Build the hierarchy of the Grouped Time Series stores = [i for i in stores_ts.columns] items = [i for i in items_ts.columns] store_items = col_names # Exogenous features mapping exog_store_items = {e: [v for v in exog_train_df.columns if v.startswith(e)] for e in store_items} exog_stores = {e:[v for v in exog_train_df.columns if v.startswith(e)] for e in stores} exog_items = {e:[v for v in exog_train_df.columns if v.find(e) != -1] for e in items} exog_total = {'total':[v for v in exog_train_df.columns if v.find('FOURIER') != -1]} # Merge all dictionaries exog_to_merge = [exog_store_items, exog_stores, exog_items, exog_total] exogenous = {k:v for x in exog_to_merge for k,v in x.items()} # Build hierarchy total = {'total': stores + items} store_h = {k: [v for v in store_items if v.startswith(k)] for k in stores} hierarchy = {**total, **store_h} # Hierarchy tree automatically created by hts ht = HierarchyTree.from_nodes(nodes=hierarchy, df=df, exogenous=exogenous) # Instanciate the auto arima model using HTSRegressor autoarima = HTSRegressor(model='auto_arima', D=1, m=7, seasonal=True, revision_method='OLS', n_jobs=12) # Fit the model to the training df that includes time series and exog_train_df # Set exogenous param to the previously built dictionary model = autoarima.fit(df, hierarchy, exogenous=exogenous) # Make predictions # Set the exogenous_df param predictions = model.predict(exogenous_df=exog_test_df, steps_ahead=365)
我还尝试过另外两种方案,在单条时间序列(如store1_item1)上验证有效:
- 遍历500条时间序列,对每条独立使用TBATS建模预测
- 遍历500条时间序列,对每条独立使用带傅里叶外生特征的auto_arima(SARIMAX)建模预测
请问上述两种方案是否合理?有没有更适合多时间序列场景的ARIMA扩展方案?
另外我想尝试用LSTM完成该任务,我是数据科学和深度学习领域新手,请问数据预处理应该采用原始长格式对门店、商品字段做独热编码,还是使用本次处理后的宽表df?
现有方案合理性判断
你提到的两种单序列独立建模的方案完全合理,适配当前场景的优势很明确:
- TBATS原生支持多周期季节性,不需要额外手动构造傅里叶外生变量,代码实现更简洁,对销量的周、年双季节性适配性好,缺点是500条序列并行训练耗时略高于SARIMAX。
- 带傅里叶项的auto_arima方案可以灵活控制季节性拟合的复杂度,你已经在单序列上验证有效,批量跑的时候只要做好异常捕获、负值截断即可落地。
针对多时间序列场景的ARIMA类扩展方案,可以选择全局ARIMA(GlobalARIMA)思路:所有序列共享相同的ARIMA阶数参数,仅拟合各自的截距和外生变量系数,相比单序列独立建模训练速度快2~3倍,预测精度损失极小,非常适合这类同分布的门店-商品销量预测场景,可以用pmdarima的批量建模接口结合自定义全局参数搜索逻辑实现。
scikit-hts预测失真的原因
你遇到的负值、预测失真问题核心是层级调和逻辑和外生变量配置冲突:
- 你构造的外生变量映射有冗余,上层节点(比如store_1、total节点)的外生变量是多个底层序列傅里叶项的拼接,维度和模型训练阶段期望的输入不匹配,即使修复了Github issue提到的代码bug,模型拟合的参数本身就是错误的。
- OLS调和方法会强制所有层级的预测结果满足聚合约束,当底层序列预测本身有偏差时,调和过程会进一步放大误差,甚至出现负值。如果坚持用scikit-hts,可以先把调和方法换成'BU'(BottomUp,底层预测后直接向上聚合),先验证底层序列预测的准确性,再考虑更复杂的调和策略。
LSTM数据预处理建议
作为新手优先选择原始长格式+独热编码的方案,优势如下:
- 长格式输入的结构更符合LSTM序列化输入的逻辑,每个样本对应
[时间步, 特征数]的维度,特征可以包含独热编码后的门店ID、商品ID、历史销量、傅里叶季节性特征、日期衍生特征(是否节假日、周几、月份等)。 - 宽表的列数会随着门店、商品数量增加暴增,输入维度太高很容易导致LSTM过拟合,训练难度大,对新手不友好。
- 长格式可以很方便地做滑动窗口构造训练样本,也支持未来扩展全局共享的特征,泛化性更好。
内容的提问来源于stack exchange,提问作者Downforu

