You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scikit-learn Pipeline中为FunctionTransformer传入多数据集/参数?

问题根源
  1. Scikit-learn原生Pipeline的设计逻辑是:仅向每个步骤的fit方法传递特征矩阵X,默认只接受处理X的转换器,不会自动传递标签y。
  2. 你使用的FunctionTransformer是无状态转换器,其包装的函数会在transform阶段执行,而transform阶段Pipeline不会传递y参数;单独测试时你手动传入了y所以能运行,但Pipeline中无法自动完成这件事。
  3. 更关键的是:采样操作需要同时修改X和y,但Scikit-learn原生Pipeline仅支持修改X的步骤,无法同步处理y。
解决方法

方法1:使用imblearn的Pipeline(最推荐,适配不平衡数据集采样场景)

imblearn库的Pipeline专门支持同时处理X和y的采样步骤,只需将你的自定义采样函数封装为采样器类:

步骤1:自定义采样器类

继承imblearn.base.BaseSampler,实现_fit_resample方法(该方法接收X和y,返回采样后的X和y):

from imblearn.base import BaseSampler

# 欠采样采样器
class UndersampleSampler(BaseSampler):
    def __init__(self, strategy_count):
        super().__init__()
        self.strategy_count = strategy_count

    def _fit_resample(self, X, y):
        # 调用你的自定义欠采样函数
        return undersample_train_set(X, y, self.strategy_count)

# SMOTE采样器
class SMOTESampler(BaseSampler):
    def __init__(self, strategy_count, cat_cols, knn):
        super().__init__()
        self.strategy_count = strategy_count
        self.cat_cols = cat_cols
        self.knn = knn

    def _fit_resample(self, X, y):
        # 调用你的自定义SMOTE采样函数
        return SMOTE_train_set(X, y, self.strategy_count, self.cat_cols, self.knn)

步骤2:构建imblearn Pipeline

from imblearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression  # 替换为你的模型

# 构建Pipeline:采样步骤 -> 模型
pipeline = Pipeline([
    ('sampling', UndersampleSampler(strategy_count={'0': 100, '1': 100})),
    ('model', LogisticRegression())
])

# 训练时自动传递X和y给采样器,同步完成X和y的采样
pipeline.fit(X_train, y_train)

方法2:自定义带状态的转换器(适配Scikit-learn原生Pipeline)

如果你坚持使用Scikit-learn原生Pipeline,需要自定义转换器类,同时注意:这种方式仅修改X,你需要手动同步处理y(或在模型训练前手动拆分采样后的X和y)。

自定义转换器类

继承BaseEstimator和TransformerMixin,在fit阶段完成采样并保存采样后的样本索引,transform阶段返回对应特征:

from sklearn.base import BaseEstimator, TransformerMixin
import numpy as np

class UndersampleTransformer(BaseEstimator, TransformerMixin):
    def __init__(self, strategy_count):
        self.strategy_count = strategy_count
        self.resampled_indices = None  # 保存采样后的样本索引
        self.resampled_y = None  # 保存采样后的标签

    def fit(self, X, y):
        X_resampled, y_resampled = undersample_train_set(X, y, self.strategy_count)
        # 保存索引(兼容DataFrame和numpy数组)
        self.resampled_indices = X_resampled.index if hasattr(X, 'index') else np.arange(len(X_resampled))
        self.resampled_y = y_resampled
        return self

    def transform(self, X):
        if hasattr(X, 'loc'):
            return X.loc[self.resampled_indices]
        else:
            return X[self.resampled_indices]

使用方式

from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression

transformer = UndersampleTransformer(strategy_count={'0': 100, '1': 100})
pipeline = Pipeline([
    ('undersample', transformer),
    ('model', LogisticRegression())
])

# 先拟合转换器,获取采样后的y
transformer.fit(X_train, y_train)
# 用采样后的X和y训练Pipeline
pipeline.fit(transformer.transform(X_train), transformer.resampled_y)
注意事项
  • 采样仅应在训练集上执行,测试集禁止采样,避免数据泄露。
  • imblearn的Pipeline会自动处理采样后的X和y传递,无需手动拆分,是处理不平衡数据集采样的最优方案。

内容的提问来源于stack exchange,提问作者baharak Al

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 12:35:38