You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scikit-learn Pipeline编码后标签未转换的问题排查

问题原因分析
  1. 打印的y_train是原始数据:你执行train_test_split后得到的y_train是直接从原始数据集分割出来的,根本没经过你创建的Pipeline处理,所以自然还是文本标签。Pipeline的预处理逻辑只有在调用fit()或transform()时才会生效,不会自动修改你的原始变量。
  2. 错误将LabelEncoder放入ColumnTransformer:ColumnTransformer的作用是处理特征矩阵X,它的输入是特征数据,不能用来处理标签y。你在preprocessor里添加处理标签列的步骤完全无效,因为训练时Pipeline只会把X传给preprocessor,不会传入y。
解决方法

方法一:单独编码标签(简单直接)

先对标签列进行Label编码,再分割数据集并训练模型:

from sklearn.preprocessing import LabelEncoder

df = self._dataset.as_dataframe()
train_feature = df[self._train_configs.feature_cols]
train_label = df[self._train_configs.target_col]

# 单独编码标签
le = LabelEncoder()
encoded_train_label = le.fit_transform(train_label)

# 创建只处理特征的Pipeline
self._model = self.create_pipeline(train_feature, self._train_configs.encoding_method, self._model)

# 分割数据集(注意用编码后的标签)
X_train, X_test, y_train, y_test = train_test_split(train_feature, encoded_train_label, random_state=0, train_size=0.8)

# 训练模型
self._model.fit(X_train, y_train)

# 此时打印y_train就是编码后的数值
print(y_train)

同时修改create_pipeline函数,移除所有处理标签的逻辑:

def create_pipeline(self, train_feature, encoding_method, model):
    if encoding_method == EncodingMethod.ONE_HOT:
        categorical_cols = [col for col in train_feature.columns if train_feature[col].dtype == 'object'] 
        categorical_transformer = Pipeline(steps=[
            ('onehot', OneHotEncoder(handle_unknown='ignore', sparse=False))
        ])
        preprocessor = ColumnTransformer(
            transformers=[('category', categorical_transformer, categorical_cols)],
            remainder='passthrough'
        )
    elif encoding_method == EncodingMethod.LABEL:
        categorical_cols = [col for col in train_feature.columns if train_feature[col].dtype == 'object']
        categorical_transformer = Pipeline(steps=[
            ('label', LabelEncoder())
        ])
        preprocessor = ColumnTransformer(
            transformers=[('category', categorical_transformer, categorical_cols)],
            remainder='passthrough'
        )
    pipeline = Pipeline(steps=[('preprocessor', preprocessor), ('classifier', model)])
    return pipeline

方法二:将标签编码整合到Pipeline(更规范)

如果想把标签编码也纳入Pipeline流程,可以使用TransformedTargetRegressor(分类场景同样适用):

from sklearn.compose import TransformedTargetRegressor

# 先创建处理特征的Pipeline(同方法一中修改后的create_pipeline)
feature_pipeline = self.create_pipeline(train_feature, self._train_configs.encoding_method, self._model)

# 包装标签编码逻辑
self._model = TransformedTargetRegressor(
    regressor=feature_pipeline,
    transformer=LabelEncoder(),
    check_inverse=False  # 分类场景不需要逆变换,设为False
)

# 分割原始数据集
X_train, X_test, y_train, y_test = train_test_split(train_feature, train_label, random_state=0, train_size=0.8)

# 训练时自动编码标签
self._model.fit(X_train, y_train)

# 查看编码后的标签可以手动用LabelEncoder转换
le = LabelEncoder()
print(le.fit_transform(y_train))

额外注意点

  • 部分Scikit-learn分类模型可直接接受字符串标签,但编码后更规范,也能避免潜在兼容问题。
  • 处理测试集标签时,必须使用训练集拟合好的LabelEncoder执行transform(),不能重新fit_transform(),防止数据泄露。

内容的提问来源于stack exchange,提问作者Stackie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 04:23:13