如何在sklearn的ColumnTransformer/Pipeline中排除指定列(含OHE生成列)?
问题根源
你之前的自定义转换器失败,核心原因有两个:
ColumnTransformer输出的是numpy数组而非DataFrame,数组没有drop方法;- 你传入的是原始列名(如"Example"),但转换后已变成OHE生成的列名(如"Example_B"),自然找不到对应列。
解决方案
方案1:在OneHotEncoder阶段直接排除目标类别(最简洁)
从源头避免生成不需要的列,直接在OHE配置中指定要丢弃的类别:
# 修改OHE管道,指定要丢弃的类别 SimpImpConstNoAns_OHE = Pipeline([ ('SimpleImputer', SimpleImputer(strategy="constant", fill_value='no_answer')), # 直接指定要丢弃的类别,这里以排除Example_B为例 ('OHE', OneHotEncoder(sparse=False, drop=['B'], categories='auto')) ]) # 后续ColumnTransformer和管道逻辑保持不变 preprocessor_transformer = ColumnTransformer([ ('pipeline-1', SimpImpMean_MinMaxScaler, ['Num_col']), ('pipeline-2', SimpImpConstNoAns_OHE, ['Example']) ], remainder='drop', verbose_feature_names_out=False) # 测试输出 result_df = pd.DataFrame(preprocessor_transformer.fit_transform(df_foo), columns=preprocessor_transformer.get_feature_names_out()) print(result_df)
方案2:自定义转换器按索引删除数组列
如果必须在转换后删除,先确定目标列的索引,再对numpy数组进行切片操作:
import numpy as np # 1. 先拟合预处理管道,获取输出列名并定位目标列索引 preprocessor_transformer.fit(df_foo) feature_names = preprocessor_transformer.get_feature_names_out() # 找到要删除的列的索引,比如排除Example_B drop_idx = [i for i, name in enumerate(feature_names) if name == 'Example_B'] # 2. 编写适配numpy数组的自定义删除器 class ColumnDropperTransformer: def __init__(self, drop_indices): self.drop_indices = drop_indices def fit(self, X, y=None): return self def transform(self, X, y=None): return np.delete(X, self.drop_indices, axis=1) # 3. 构建完整管道 from sklearn.pipeline import make_pipeline full_pipeline = make_pipeline(preprocessor_transformer, ColumnDropperTransformer(drop_idx)) # 测试输出 result = full_pipeline.fit_transform(df_foo) result_df = pd.DataFrame(result, columns=[name for name in feature_names if name != 'Example_B']) print(result_df)
方案3:用FunctionTransformer转DataFrame按列名删除
如果更习惯用列名操作,可借助FunctionTransformer将数组转成DataFrame处理后再转回数组:
from sklearn.preprocessing import FunctionTransformer def drop_target_columns(X, feature_names, drop_cols): # 数组转DataFrame df = pd.DataFrame(X, columns=feature_names) # 删除指定列 df = df.drop(drop_cols, axis=1) # 转回numpy数组 return df.values # 1. 拟合预处理管道,获取输出列名 preprocessor_transformer.fit(df_foo) feature_names = preprocessor_transformer.get_feature_names_out() # 指定要删除的列 target_drop_cols = ['Example_B'] # 2. 创建转换器 col_dropper = FunctionTransformer( drop_target_columns, kw_args={'feature_names': feature_names, 'drop_cols': target_drop_cols} ) # 3. 构建完整管道 full_pipeline = make_pipeline(preprocessor_transformer, col_dropper) # 测试输出 result = full_pipeline.fit_transform(df_foo) result_df = pd.DataFrame(result, columns=[name for name in feature_names if name not in target_drop_cols]) print(result_df)
内容的提问来源于stack exchange,提问作者chvieira2
相关产品推荐
相关产品推荐

