You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用sklearn pipeline实现标签预处理?运行报错如何解决?

报错根因

你定义的preprocessor_pipeline是针对多列特征设计的ColumnTransformer组件,要求输入必须是二维表格(形状为(样本数, 特征数)),但你传入的y_train是一维的pandas Series(形状仅为(样本数,)),代码尝试访问shape[1]获取列数时自然触发索引越界。

核心原理解释
  • 特征与标签的预处理是完全独立的两个流程,不可共用同一个处理管道:特征是多维度异构数据,同时包含数值、分类等不同类型的列,因此需要ColumnTransformer实现分列差异化处理;标签是单维度数据,回归任务下为单数值列,分类任务下为单类别列,无需多列并行处理逻辑,单独做针对性处理即可。
  • sklearn所有内置变换器的默认输入要求为二维数组,一维的Series/数组需要先通过.values.reshape(-1, 1)升维后才能正常传入处理。
  • 标签预处理要严格遵循「训练集拟合、训练/测试集分别转换」的规则,不可使用测试集的统计量处理标签,避免数据泄露。
修正后的完整代码

以下代码移除了对标签调用特征处理管道的错误逻辑,同时新增了单独的标签预处理流程(如果标签无缺失值、且后续使用树类模型训练,可直接跳过标签预处理步骤,返回原始的y_train、y_test即可):

# Data Preprocessing
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from icecream import ic

def diamond_preprocess(data_dir):
    data = pd.read_csv(data_dir)
    cleaned_data = data.drop(['id', 'depth_percent'], axis=1)  # 移除不需要的特征

    x = cleaned_data.drop(['price'], axis=1)  # 特征集
    y = cleaned_data['price']  # 标签列

    x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=99)

    # ---------------------- 特征预处理流程 ----------------------
    numerical_features = x_train.select_dtypes(include=['int64', 'float64']).columns.tolist()
    categorical_features = x_train.select_dtypes(include=['object']).columns.tolist()

    numerical_transformer = Pipeline(steps=[
        ('imputer', SimpleImputer(strategy='median')),  # 数值列缺失值用中位数填充
        ('scaler', StandardScaler())  # 数值列标准化
    ])

    categorical_transformer = Pipeline(steps=[
        ('imputer', SimpleImputer(strategy='constant', fill_value='missing')),  # 分类列缺失值填充为'missing'
        ('onehot', OneHotEncoder(handle_unknown='ignore'))  # 分类列独热编码
    ])

    preprocessor_pipeline = ColumnTransformer(
        transformers=[
            ('num', numerical_transformer, numerical_features),
            ('cat', categorical_transformer, categorical_features)
        ])

    # 仅用训练集特征拟合预处理管道,再转换训练/测试集特征
    preprocessor_pipeline.fit(x_train)
    x_train = pd.DataFrame(preprocessor_pipeline.transform(x_train))
    x_test = pd.DataFrame(preprocessor_pipeline.transform(x_test))

    # ---------------------- 标签预处理流程 ----------------------
    # 仅当标签有缺失值、或后续使用对标签量级敏感的模型(线性回归、神经网络等)时需要执行
    y_transformer = Pipeline(steps=[
        ('imputer', SimpleImputer(strategy='median')),
        ('scaler', StandardScaler())
    ])
    # 一维标签升维后拟合、转换
    y_transformer.fit(y_train.values.reshape(-1, 1))
    y_train = pd.DataFrame(y_transformer.transform(y_train.values.reshape(-1, 1)))
    y_test = pd.DataFrame(y_transformer.transform(y_test.values.reshape(-1, 1)))

    # 后续需要把预测值还原为原始价格量级时,可调用:y_transformer.inverse_transform(y_pred.reshape(-1,1))

    return x_train, x_test, y_train, y_test

内容的提问来源于stack exchange,提问作者Luleo_Primoc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 20:15:00