You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建含随机过采样与自定义离群值移除器的Pipeline时遇样本数不匹配错误

机器学习Pipeline中过采样+离群值移除的样本数不匹配问题解决

问题

构建包含随机过采样(RandomOverSampler)和自定义离群值移除器(OutlierRemover)的机器学习Pipeline时,出现如下错误:

ValueError: Found input variables with inconsistent numbers of samples: [310231, 363920]

本质原因是离群值移除后X的样本数减少,但y未同步更新,导致后续模型训练时样本数不匹配。

核心原因

  1. scikit-learn原生的make_pipeline不支持转换器(Transformer)在transform方法中同时返回X和y,只能返回X,因此自定义的OutlierRemover生成的新y无法被Pipeline传递给后续模型。
  2. 原OutlierRemover的fit方法逻辑错误:直接保存训练数据的X和y,而非学习训练数据的统计量(如离群值上下界),既不符合scikit-learn API规范,也容易导致数据泄露。

解决方案

修改步骤

  1. 使用imblearn库的Pipeline替代scikit-learn原生make_pipeline:imblearn的Pipeline支持处理同时返回X和y的转换器(如过采样器、自定义离群值移除器)。
  2. 修正自定义OutlierRemover的逻辑:在fit阶段学习训练数据的离群值上下界,transform阶段使用这些预计算值过滤样本,确保X和y同步更新。
  3. 调整最佳估计器的获取方式:通过Pipeline的named_steps属性获取内部的RandomizedSearchCV实例,避免直接引用外部变量导致的状态混乱。

修改后的完整代码

修正后的自定义离群值移除器

import numpy as np
from sklearn.base import BaseEstimator, TransformerMixin

class OutlierRemover(BaseEstimator, TransformerMixin):
    def __init__(self, columns):
        self.columns = columns
        self.bounds = {}  # 存储每个列的离群值上下界

    def fit(self, X, y=None):
        # 仅在训练数据上计算各列的离群值边界
        for col in self.columns:
            q25, q75 = np.percentile(X[col], 25), np.percentile(X[col], 75)
            iqr = q75 - q25
            cut_off = iqr * 1.5
            lower, upper = q25 - cut_off, q75 + cut_off
            self.bounds[col] = (lower, upper)
        return self

    def transform(self, X, y=None):
        new_X = X.copy()
        # 收集所有需要保留的样本索引
        indices_to_keep = new_X.index
        for col in self.columns:
            lower, upper = self.bounds[col]
            # 保留当前列在范围内的样本索引
            indices_to_keep = indices_to_keep.intersection(
                new_X[(new_X[col] >= lower) & (new_X[col] <= upper)].index
            )
        # 过滤X
        new_X = new_X.loc[indices_to_keep]
        # 同步过滤y(如果存在)
        if y is not None:
            new_y = y.loc[indices_to_keep]
            return new_X, new_y
        else:
            return new_X

修正后的Pipeline与交叉验证代码

from imblearn.pipeline import Pipeline
from imblearn.over_sampling import RandomOverSampler
from sklearn.model_selection import StratifiedKFold, RandomizedSearchCV
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
import numpy as np

accuracy_lst = []
precision_lst = []
recall_lst = []
f1_lst = []
auc_lst = []

kf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
log_reg_params = {"penalty": ['l1', 'l2'], 'C': [0.001, 0.01, 0.1, 1, 10, 100, 1000]}
rand_log_reg = RandomizedSearchCV(LogisticRegression(max_iter=200), log_reg_params, n_iter=4)

for train, test in kf.split(Org_X_train, Org_y_train):
    X_train, X_test = Org_X_train.iloc[train], Org_X_train.iloc[test]
    y_train, y_test = Org_y_train.iloc[train], Org_y_train.iloc[test]
    
    # 使用imblearn的Pipeline定义流程
    pipeline = Pipeline([
        ('oversampler', RandomOverSampler(random_state=42)),
        ('outlier_remover', OutlierRemover(columns=['V14', 'V12', 'V10', 'V4', 'V11', 'V2'])),
        ('random_search', rand_log_reg)
    ])

    model = pipeline.fit(X_train, y_train)
    # 从Pipeline中获取训练后的RandomizedSearchCV实例,再取最佳估计器
    best_est = model.named_steps['random_search'].best_estimator_
    prediction = best_est.predict(X_test)

    # 计算并保存各指标
    accuracy_lst.append(accuracy_score(y_test, prediction))
    precision_lst.append(precision_score(y_test, prediction))
    recall_lst.append(recall_score(y_test, prediction))
    f1_lst.append(f1_score(y_test, prediction))
    auc_lst.append(roc_auc_score(y_test, prediction))

# 输出平均指标
print("Accuracy:", np.mean(accuracy_lst))
print("Precision:", np.mean(precision_lst))
print("Recall:", np.mean(recall_lst))
print("F1 Score:",  np.mean(f1_lst))
print("AUC Score:",  np.mean(auc_lst))

内容的提问来源于stack exchange,提问作者Roterun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 02:15:15