You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Rocket变换仅输出1行结果的技术求助

问题原因

sktime的Rocket变换是为面板时间序列数据设计的,它要求输入数据符合时间序列样本结构:每个样本对应一段时间序列(单变量或多变量均可)。

你当前的二维DataFrame被Rocket错误解析成了1个样本——整个DataFrame被当作是1个包含5个变量、长度为1000的时间序列,因此变换后仅输出1行结果。

修正方案

需要将表格数据转换为sktime要求的面板时间序列格式,以下是两种常用实现方式:

方式1:转换为三维numpy数组

将原数据重塑为(n_samples, n_dimensions, n_timesteps)的结构。如果你的每个样本是单变量时间序列(含5个时间步),可以这样处理:

import pandas as pd
from sktime.transformations.panel.rocket import Rocket
from sklearn.datasets import make_classification

X, y = make_classification(
    n_samples=1000,
    n_features=5,
    n_informative=3,
    n_classes=2,
    random_state=999
)

# 重塑为三维数组:(1000个样本, 1个变量, 5个时间步)
X_rocket = X.reshape(1000, 1, 5)

rocket = Rocket(num_kernels=1000, random_state=0)
rocket.fit(X_rocket)
X_train_transform = rocket.transform(X_rocket)

print(X_train_transform.shape)  # 输出 (1000, 2000)

方式2:使用sktime的Panel DataFrame格式

通过多层索引DataFrame表示面板数据,第一层为样本ID,第二层为时间步:

import pandas as pd
from sktime.transformations.panel.rocket import Rocket
from sklearn.datasets import make_classification

X, y = make_classification(
    n_samples=1000,
    n_features=5,
    n_informative=3,
    n_classes=2,
    random_state=999
)

# 构建多层索引面板数据
df_list = []
for idx in range(X.shape[0]):
    sample_df = pd.DataFrame(X[idx], columns=["value"])
    sample_df["sample_id"] = idx
    sample_df["time_step"] = range(5)
    df_list.append(sample_df)

panel_df = pd.concat(df_list).set_index(["sample_id", "time_step"])

rocket = Rocket(num_kernels=1000, random_state=0)
rocket.fit(panel_df)
X_train_transform = rocket.transform(panel_df)

print(X_train_transform.shape)  # 输出 (1000, 2000)
额外说明

如果你的数据本质是传统表格数据(非时间序列),Rocket并非最优选择。若确实需要用Rocket处理,必须保证输入符合它的时间序列样本逻辑——每个样本对应一段时间序列,而非一行特征值。

内容的提问来源于stack exchange,提问作者Saetthakij Naothaworn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 10:20:11