如何用sklearn pipeline实现标签预处理?运行报错如何解决?
报错根因
你定义的preprocessor_pipeline是针对多列特征设计的ColumnTransformer组件,要求输入必须是二维表格(形状为(样本数, 特征数)),但你传入的y_train是一维的pandas Series(形状仅为(样本数,)),代码尝试访问shape[1]获取列数时自然触发索引越界。
核心原理解释
- 特征与标签的预处理是完全独立的两个流程,不可共用同一个处理管道:特征是多维度异构数据,同时包含数值、分类等不同类型的列,因此需要
ColumnTransformer实现分列差异化处理;标签是单维度数据,回归任务下为单数值列,分类任务下为单类别列,无需多列并行处理逻辑,单独做针对性处理即可。 - sklearn所有内置变换器的默认输入要求为二维数组,一维的Series/数组需要先通过
.values.reshape(-1, 1)升维后才能正常传入处理。 - 标签预处理要严格遵循「训练集拟合、训练/测试集分别转换」的规则,不可使用测试集的统计量处理标签,避免数据泄露。
修正后的完整代码
以下代码移除了对标签调用特征处理管道的错误逻辑,同时新增了单独的标签预处理流程(如果标签无缺失值、且后续使用树类模型训练,可直接跳过标签预处理步骤,返回原始的y_train、y_test即可):
# Data Preprocessing import pandas as pd from sklearn.compose import ColumnTransformer from sklearn.impute import SimpleImputer from sklearn.model_selection import train_test_split from sklearn.pipeline import Pipeline from sklearn.preprocessing import StandardScaler, OneHotEncoder from icecream import ic def diamond_preprocess(data_dir): data = pd.read_csv(data_dir) cleaned_data = data.drop(['id', 'depth_percent'], axis=1) # 移除不需要的特征 x = cleaned_data.drop(['price'], axis=1) # 特征集 y = cleaned_data['price'] # 标签列 x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=99) # ---------------------- 特征预处理流程 ---------------------- numerical_features = x_train.select_dtypes(include=['int64', 'float64']).columns.tolist() categorical_features = x_train.select_dtypes(include=['object']).columns.tolist() numerical_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='median')), # 数值列缺失值用中位数填充 ('scaler', StandardScaler()) # 数值列标准化 ]) categorical_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='constant', fill_value='missing')), # 分类列缺失值填充为'missing' ('onehot', OneHotEncoder(handle_unknown='ignore')) # 分类列独热编码 ]) preprocessor_pipeline = ColumnTransformer( transformers=[ ('num', numerical_transformer, numerical_features), ('cat', categorical_transformer, categorical_features) ]) # 仅用训练集特征拟合预处理管道,再转换训练/测试集特征 preprocessor_pipeline.fit(x_train) x_train = pd.DataFrame(preprocessor_pipeline.transform(x_train)) x_test = pd.DataFrame(preprocessor_pipeline.transform(x_test)) # ---------------------- 标签预处理流程 ---------------------- # 仅当标签有缺失值、或后续使用对标签量级敏感的模型(线性回归、神经网络等)时需要执行 y_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='median')), ('scaler', StandardScaler()) ]) # 一维标签升维后拟合、转换 y_transformer.fit(y_train.values.reshape(-1, 1)) y_train = pd.DataFrame(y_transformer.transform(y_train.values.reshape(-1, 1))) y_test = pd.DataFrame(y_transformer.transform(y_test.values.reshape(-1, 1))) # 后续需要把预测值还原为原始价格量级时,可调用:y_transformer.inverse_transform(y_pred.reshape(-1,1)) return x_train, x_test, y_train, y_test
内容的提问来源于stack exchange,提问作者Luleo_Primoc
相关产品推荐
相关产品推荐

