You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sktime TimeSeriesForestRegressor遇数据格式错误的解决方法

解决sktime TimeSeriesForestRegressor的数据格式错误问题

错误原因

错误提示明确指出X不符合sktime对Panel(面板时间序列)数据的要求。TimeSeriesForestRegressor是针对时间序列回归的模型,它要求输入的X是每个样本为一段时间序列的面板数据,而非普通的二维表格DataFrame。

解决方案

以下两种方式可将数据转换为sktime支持的格式:

方案1:使用MultiIndex DataFrame(sktime推荐格式)

这种格式通过双层索引区分不同样本和时间步,结构清晰:

import numpy as np
import pandas as pd
from sktime.regression.interval_based import TimeSeriesForestRegressor

# 生成原始数据
rand = np.random.random((200, 3))
# 定义每个样本的时间步长(示例设为2,即200行=100个样本×2个时间步,可按需调整)
n_timepoints = 2
n_instances = len(rand) // n_timepoints

# 构造双层索引:第一层为样本ID,第二层为时间步
instance_ids = np.repeat(range(n_instances), n_timepoints)
time_ids = np.tile(range(n_timepoints), n_instances)
multi_index = pd.MultiIndex.from_arrays([instance_ids, time_ids], names=["instance", "time"])

# 转换特征数据为MultiIndex格式
X = pd.DataFrame(rand[:, :2], index=multi_index, columns=["feature_0", "feature_1"])
# 每个样本对应一个目标值(示例取每个样本第一个时间步的目标值,可按需调整聚合方式)
y = pd.Series(rand[::n_timepoints, 2], index=range(n_instances))

# 训练模型
forecaster = TimeSeriesForestRegressor()
forecaster.fit(X=X, y=y)

方案2:使用3D numpy数组

直接将数据重塑为(样本数, 特征数, 时间步数)的3D数组:

import numpy as np
import pandas as pd
from sktime.regression.interval_based import TimeSeriesForestRegressor

rand = np.random.random((200, 3))
n_timepoints = 2
n_instances = len(rand) // n_timepoints

# 重塑特征数据为3D数组
X = rand[:, :2].reshape(n_instances, 2, n_timepoints)
# 目标变量每个样本对应一个值
y = rand[::n_timepoints, 2]

forecaster = TimeSeriesForestRegressor()
forecaster.fit(X=X, y=y)

额外提示

如果你的需求是普通表格回归(每行是独立样本,特征为静态值而非时间序列),sktime的时间序列回归器并不适用,应使用scikit-learn的常规回归模型(如RandomForestRegressor)。

内容的提问来源于stack exchange,提问作者Mert Arda Asar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 06:05:50