You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在sktime中递归使用WindowSummarizer进行未来期预测?

问题分析与解决方案

报错原因

你的代码出现NaN错误,核心问题有三个:

  1. 窗口方向写反:WindowSummarizer的lag_feature中,mean/std等统计量的窗口参数(如[0,7])是正向偏移,意味着要引用未来7步的数据,但预测时未来数据不存在,直接生成NaN。
  2. 特征包含当前目标值:lag设置了0,把当前要预测的目标值作为特征,训练时属于数据泄露,预测时无法获取该值,必然产生NaN。
  3. 递归窗口长度未明确:make_reduction的window_length设为None,无法匹配你的最大窗口需求(28步),导致特征生成时数据不足。

修正后的代码

#!pip install sktime[all_extras]

import pandas as pd
import numpy as np
from sktime.transformations.series.summarize import WindowSummarizer
from sktime.forecasting.compose import ForecastingPipeline
from sklearn.linear_model import LinearRegression
from sktime.datasets import load_airline
from sktime.forecasting.compose import make_reduction

y = load_airline()

# 修正1:窗口参数改为负向偏移,引用历史数据;移除lag=0
kwargs = {
        "lag_feature": {
            "lag": [1, 2],  # 用前1步、前2步的值作为特征
            "mean": [[-7, 0], [-14, 0], [-28, 0]],  # 过去7/14/28步到当前的均值
            "std": [[-7, 0], [-14, 0], [-28, 0]],
            "kurt": [[-7, 0], [-14, 0], [-28, 0]],
            "skew": [[-7, 0], [-14, 0], [-28, 0]]
        }
    }

# 修正2:明确window_length为最大窗口长度+1(因为[-28,0]包含29个点)
forecaster = make_reduction(
    LinearRegression(),
    scitype="tabular-regressor",
    transformers=[WindowSummarizer(**kwargs, n_jobs=1, truncate="bfill")],
    window_length=29,  # 匹配最大窗口[-28,0]的长度
    strategy="recursive"
)

pipe = ForecastingPipeline(
    steps=[
        ("forecaster", forecaster),
    ]
)

model = pipe.fit(y)

# 测试远期预测
predictions = model.predict(fh=np.arange(1, 100, 1))
print(predictions)

关键修改说明

  • 窗口方向修正:将[0,7]改为[-7,0],表示计算从当前时刻往前7步到当前的窗口统计量,确保只用已有历史数据(包括递归预测出的结果)。
  • 移除lag=0:只保留过去的滞后值作为特征,避免数据泄露和预测时的NaN问题。
  • 设置window_length:最大窗口是[-28,0],共29个数据点,因此window_length设为29,确保递归时每次都有足够的数据生成特征。

内容的提问来源于stack exchange,提问作者Mairon Chaves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 04:39:25