如何在Scikit-learn Pipeline中为FunctionTransformer传入多数据集/参数?
问题根源
- Scikit-learn原生
Pipeline的设计逻辑是:仅向每个步骤的fit方法传递特征矩阵X,默认只接受处理X的转换器,不会自动传递标签y。 - 你使用的
FunctionTransformer是无状态转换器,其包装的函数会在transform阶段执行,而transform阶段Pipeline不会传递y参数;单独测试时你手动传入了y所以能运行,但Pipeline中无法自动完成这件事。 - 更关键的是:采样操作需要同时修改
X和y,但Scikit-learn原生Pipeline仅支持修改X的步骤,无法同步处理y。
解决方法
方法1:使用imblearn的Pipeline(最推荐,适配不平衡数据集采样场景)
imblearn库的Pipeline专门支持同时处理X和y的采样步骤,只需将你的自定义采样函数封装为采样器类:
步骤1:自定义采样器类
继承imblearn.base.BaseSampler,实现_fit_resample方法(该方法接收X和y,返回采样后的X和y):
from imblearn.base import BaseSampler # 欠采样采样器 class UndersampleSampler(BaseSampler): def __init__(self, strategy_count): super().__init__() self.strategy_count = strategy_count def _fit_resample(self, X, y): # 调用你的自定义欠采样函数 return undersample_train_set(X, y, self.strategy_count) # SMOTE采样器 class SMOTESampler(BaseSampler): def __init__(self, strategy_count, cat_cols, knn): super().__init__() self.strategy_count = strategy_count self.cat_cols = cat_cols self.knn = knn def _fit_resample(self, X, y): # 调用你的自定义SMOTE采样函数 return SMOTE_train_set(X, y, self.strategy_count, self.cat_cols, self.knn)
步骤2:构建imblearn Pipeline
from imblearn.pipeline import Pipeline from sklearn.linear_model import LogisticRegression # 替换为你的模型 # 构建Pipeline:采样步骤 -> 模型 pipeline = Pipeline([ ('sampling', UndersampleSampler(strategy_count={'0': 100, '1': 100})), ('model', LogisticRegression()) ]) # 训练时自动传递X和y给采样器,同步完成X和y的采样 pipeline.fit(X_train, y_train)
方法2:自定义带状态的转换器(适配Scikit-learn原生Pipeline)
如果你坚持使用Scikit-learn原生Pipeline,需要自定义转换器类,同时注意:这种方式仅修改X,你需要手动同步处理y(或在模型训练前手动拆分采样后的X和y)。
自定义转换器类
继承BaseEstimator和TransformerMixin,在fit阶段完成采样并保存采样后的样本索引,transform阶段返回对应特征:
from sklearn.base import BaseEstimator, TransformerMixin import numpy as np class UndersampleTransformer(BaseEstimator, TransformerMixin): def __init__(self, strategy_count): self.strategy_count = strategy_count self.resampled_indices = None # 保存采样后的样本索引 self.resampled_y = None # 保存采样后的标签 def fit(self, X, y): X_resampled, y_resampled = undersample_train_set(X, y, self.strategy_count) # 保存索引(兼容DataFrame和numpy数组) self.resampled_indices = X_resampled.index if hasattr(X, 'index') else np.arange(len(X_resampled)) self.resampled_y = y_resampled return self def transform(self, X): if hasattr(X, 'loc'): return X.loc[self.resampled_indices] else: return X[self.resampled_indices]
使用方式
from sklearn.pipeline import Pipeline from sklearn.linear_model import LogisticRegression transformer = UndersampleTransformer(strategy_count={'0': 100, '1': 100}) pipeline = Pipeline([ ('undersample', transformer), ('model', LogisticRegression()) ]) # 先拟合转换器,获取采样后的y transformer.fit(X_train, y_train) # 用采样后的X和y训练Pipeline pipeline.fit(transformer.transform(X_train), transformer.resampled_y)
注意事项
- 采样仅应在训练集上执行,测试集禁止采样,避免数据泄露。
- imblearn的Pipeline会自动处理采样后的
X和y传递,无需手动拆分,是处理不平衡数据集采样的最优方案。
内容的提问来源于stack exchange,提问作者baharak Al
相关产品推荐
相关产品推荐

