为何ARIMA模型预测结果为直线?基于PEMS08数据集的交通预测问题
ARIMA模型预测交通流量异常排查
问题描述
我基于为ASTGCN预处理的PEMS08数据集,搭建ARIMA模型预测单个节点的道路交通流量。使用statsmodels结合pmdarima的auto_arima自动筛选最优ARIMA参数,但测试集预测结果是一条无波动的直线,完全没能还原原时间序列的变化趋势。
代码实现
import numpy as np import pandas as pd import matplotlib.pyplot as plt from pmdarima import auto_arima from statsmodels.tsa.arima.model import ARIMA from sklearn.metrics import mean_absolute_error, mean_squared_error # 加载为ASTGCN预处理的数据集 data = np.load('data/PEMS08/PEMS08_r1_d0_w0_astcgn.npz') # 提取节点0的时间序列数据 data_node_0 = data['train_target'][:, 0, :] # (样本数, 预测步长) # 将数据重组为单维时间序列 time_series_node_0 = data_node_0.flatten() # 划分训练集与测试集 train_size = int(len(time_series_node_0) * 0.8) train, test = time_series_node_0[:train_size], time_series_node_0[train_size:] # 使用auto_arima确定ARIMA最优参数 stepwise_model = auto_arima(train, seasonal=False, trace=True, error_action='ignore', suppress_warnings=True) print(stepwise_model.summary()) # 基于最优参数构建ARIMA模型 tuned_order = stepwise_model.order model = ARIMA(train, order=tuned_order) model_fit = model.fit() # 测试集预测 start = len(train) end = len(train) + len(test) - 1 predictions = model_fit.predict(start=start, end=end, typ='levels') # 模型性能评估 mae = mean_absolute_error(test, predictions) mse = mean_squared_error(test, predictions) rmse = np.sqrt(mse) mape = np.mean(np.abs((test - predictions) / test)) * 100 print(f'MAE: {mae:.4f}') print(f'RMSE: {rmse:.4f}') print(f'MAPE: {mape:.4f}%') # 绘制预测结果与真实值对比图 plt.figure(figsize=(10, 5)) plt.plot(test, label='真实值', color='blue') plt.plot(predictions, label='ARIMA预测值', color='red') plt.title('节点0的ARIMA预测值与真实值对比') plt.xlabel('时间步') plt.ylabel('流量值') plt.legend() plt.grid() plt.show()
排查方向
- 检查数据集结构:确认
train_target的维度是否符合预期,扁平化操作是否破坏了时间序列的连续性 - 验证auto_arima参数:是否因
seasonal=False忽略了交通流量的日/周季节性规律,导致模型无法捕捉波动 - 查看模型参数结果:检查auto_arima输出的最优阶数是否为(0,0,0),这种情况下模型会输出常数预测值
- 数据平稳性检验:ARIMA要求序列平稳,若原始序列非平稳,需重新调整差分阶数(d值)
内容的提问来源于stack exchange,提问作者Vincenzo Pallini
相关产品推荐
相关产品推荐

