You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将多列训练的ARIMA模型合并以实现时间序列预测?

问题描述

我有一个包含100行、1000+列的时间序列DataFrame,所有列相互独立。我需要针对每一列运行ARIMA模型,但当前代码遍历列训练时会覆盖之前的模型,仅保留最后一列的模型用于测试集预测,导致预测结果过拟合。

我希望将所有训练好的模型的学习结果合并(比如取各模型预测值的平均值),得到一个统一的预测结果用于测试集。以下是示例数据:

date                        Col 1     Col 2     Col 3      Col 4
2001-07-21 10:00:00+05:00    45          51       31         3  
2001-07-21 10:15:00+05:00    46          50       32         3
2001-07-21 10:30:00+05:00    47          51       34         7
2001-07-21 10:45:00+05:00    50          50       33         9
2001-07-21 11:00:00+05:00    55          51       32         8
2001-07-21 11:15:00+05:00    52          73       34         11
2001-07-21 11:30:00+05:00    51          72       30         14

我的错误代码如下:

# training set includes all columns except the last and test set includes only last column.
train = df.iloc[:, :-1]
test = df.iloc[:,-1:]

order = (1,2,3) 

for col in train.columns:
  model = ARIMA(train[col], order = order)  # training every column in training set
  model = model.fit()
model.summary()

predictions = model.predict(len(test))
解决方案

1. 保存所有训练好的ARIMA模型

用字典存储每一列对应的训练好的模型,避免覆盖。字典的键为列名,值为对应列训练完成的ARIMA模型实例。

2. 生成所有模型的预测值并合并

遍历字典中的每个模型,生成对应测试集长度的预测值,再将所有预测值取平均得到最终的合并预测结果。

修正后的代码:

from statsmodels.tsa.arima.model import ARIMA
import pandas as pd

train = df.iloc[:, :-1]
test = df.iloc[:, -1:]
order = (1, 2, 3)

# 存储所有训练好的模型
models = {}
for col in train.columns:
    model = ARIMA(train[col], order=order)
    fitted_model = model.fit()
    models[col] = fitted_model

# 生成每个模型的预测值
all_predictions = []
for col, model in models.items():
    # 预测测试集长度的数值,起始点对应训练集结束位置
    pred = model.predict(start=len(train), end=len(train)+len(test)-1)
    all_predictions.append(pred)

# 将所有预测值合并为DataFrame,计算平均值作为最终预测
predictions_df = pd.concat(all_predictions, axis=1)
final_predictions = predictions_df.mean(axis=1)

说明

  • 字典models确保所有列的训练模型都被保留,不会被后续循环覆盖。
  • 预测的起始和结束位置设置为len(train)到len(train)+len(test)-1,保证预测的时间步长与测试集完全匹配。
  • 除了简单平均,你也可以根据模型拟合效果(比如AIC值、R²)设置加权平均,进一步优化合并后的预测结果。

内容的提问来源于stack exchange,提问作者A Newbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 02:50:40