You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在sklearn的ColumnTransformer/Pipeline中排除指定列(含OHE生成列)?

问题根源

你之前的自定义转换器失败,核心原因有两个:

  1. ColumnTransformer输出的是numpy数组而非DataFrame,数组没有drop方法;
  2. 你传入的是原始列名(如"Example"),但转换后已变成OHE生成的列名(如"Example_B"),自然找不到对应列。

解决方案

方案1:在OneHotEncoder阶段直接排除目标类别(最简洁)

从源头避免生成不需要的列,直接在OHE配置中指定要丢弃的类别:

# 修改OHE管道,指定要丢弃的类别
SimpImpConstNoAns_OHE = Pipeline([
    ('SimpleImputer', SimpleImputer(strategy="constant", fill_value='no_answer')),
    # 直接指定要丢弃的类别,这里以排除Example_B为例
    ('OHE', OneHotEncoder(sparse=False, drop=['B'], categories='auto'))
])

# 后续ColumnTransformer和管道逻辑保持不变
preprocessor_transformer = ColumnTransformer([
    ('pipeline-1', SimpImpMean_MinMaxScaler, ['Num_col']),
    ('pipeline-2', SimpImpConstNoAns_OHE, ['Example'])
     ],
    remainder='drop',
    verbose_feature_names_out=False)

# 测试输出
result_df = pd.DataFrame(preprocessor_transformer.fit_transform(df_foo), 
                         columns=preprocessor_transformer.get_feature_names_out())
print(result_df)

方案2:自定义转换器按索引删除数组列

如果必须在转换后删除,先确定目标列的索引,再对numpy数组进行切片操作:

import numpy as np

# 1. 先拟合预处理管道,获取输出列名并定位目标列索引
preprocessor_transformer.fit(df_foo)
feature_names = preprocessor_transformer.get_feature_names_out()
# 找到要删除的列的索引,比如排除Example_B
drop_idx = [i for i, name in enumerate(feature_names) if name == 'Example_B']

# 2. 编写适配numpy数组的自定义删除器
class ColumnDropperTransformer:
    def __init__(self, drop_indices):
        self.drop_indices = drop_indices

    def fit(self, X, y=None):
        return self

    def transform(self, X, y=None):
        return np.delete(X, self.drop_indices, axis=1)

# 3. 构建完整管道
from sklearn.pipeline import make_pipeline
full_pipeline = make_pipeline(preprocessor_transformer, ColumnDropperTransformer(drop_idx))

# 测试输出
result = full_pipeline.fit_transform(df_foo)
result_df = pd.DataFrame(result, columns=[name for name in feature_names if name != 'Example_B'])
print(result_df)

方案3:用FunctionTransformer转DataFrame按列名删除

如果更习惯用列名操作,可借助FunctionTransformer将数组转成DataFrame处理后再转回数组:

from sklearn.preprocessing import FunctionTransformer

def drop_target_columns(X, feature_names, drop_cols):
    # 数组转DataFrame
    df = pd.DataFrame(X, columns=feature_names)
    # 删除指定列
    df = df.drop(drop_cols, axis=1)
    # 转回numpy数组
    return df.values

# 1. 拟合预处理管道,获取输出列名
preprocessor_transformer.fit(df_foo)
feature_names = preprocessor_transformer.get_feature_names_out()
# 指定要删除的列
target_drop_cols = ['Example_B']

# 2. 创建转换器
col_dropper = FunctionTransformer(
    drop_target_columns, 
    kw_args={'feature_names': feature_names, 'drop_cols': target_drop_cols}
)

# 3. 构建完整管道
full_pipeline = make_pipeline(preprocessor_transformer, col_dropper)

# 测试输出
result = full_pipeline.fit_transform(df_foo)
result_df = pd.DataFrame(result, columns=[name for name in feature_names if name not in target_drop_cols])
print(result_df)

内容的提问来源于stack exchange,提问作者chvieira2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 17:40:41