构建含随机过采样与自定义离群值移除器的Pipeline时遇样本数不匹配错误
机器学习Pipeline中过采样+离群值移除的样本数不匹配问题解决
问题
构建包含随机过采样(RandomOverSampler)和自定义离群值移除器(OutlierRemover)的机器学习Pipeline时,出现如下错误:
ValueError: Found input variables with inconsistent numbers of samples: [310231, 363920]
本质原因是离群值移除后X的样本数减少,但y未同步更新,导致后续模型训练时样本数不匹配。
核心原因
- scikit-learn原生的
make_pipeline不支持转换器(Transformer)在transform方法中同时返回X和y,只能返回X,因此自定义的OutlierRemover生成的新y无法被Pipeline传递给后续模型。 - 原
OutlierRemover的fit方法逻辑错误:直接保存训练数据的X和y,而非学习训练数据的统计量(如离群值上下界),既不符合scikit-learn API规范,也容易导致数据泄露。
解决方案
修改步骤
- 使用
imblearn库的Pipeline替代scikit-learn原生make_pipeline:imblearn的Pipeline支持处理同时返回X和y的转换器(如过采样器、自定义离群值移除器)。 - 修正自定义
OutlierRemover的逻辑:在fit阶段学习训练数据的离群值上下界,transform阶段使用这些预计算值过滤样本,确保X和y同步更新。 - 调整最佳估计器的获取方式:通过Pipeline的
named_steps属性获取内部的RandomizedSearchCV实例,避免直接引用外部变量导致的状态混乱。
修改后的完整代码
修正后的自定义离群值移除器
import numpy as np from sklearn.base import BaseEstimator, TransformerMixin class OutlierRemover(BaseEstimator, TransformerMixin): def __init__(self, columns): self.columns = columns self.bounds = {} # 存储每个列的离群值上下界 def fit(self, X, y=None): # 仅在训练数据上计算各列的离群值边界 for col in self.columns: q25, q75 = np.percentile(X[col], 25), np.percentile(X[col], 75) iqr = q75 - q25 cut_off = iqr * 1.5 lower, upper = q25 - cut_off, q75 + cut_off self.bounds[col] = (lower, upper) return self def transform(self, X, y=None): new_X = X.copy() # 收集所有需要保留的样本索引 indices_to_keep = new_X.index for col in self.columns: lower, upper = self.bounds[col] # 保留当前列在范围内的样本索引 indices_to_keep = indices_to_keep.intersection( new_X[(new_X[col] >= lower) & (new_X[col] <= upper)].index ) # 过滤X new_X = new_X.loc[indices_to_keep] # 同步过滤y(如果存在) if y is not None: new_y = y.loc[indices_to_keep] return new_X, new_y else: return new_X
修正后的Pipeline与交叉验证代码
from imblearn.pipeline import Pipeline from imblearn.over_sampling import RandomOverSampler from sklearn.model_selection import StratifiedKFold, RandomizedSearchCV from sklearn.linear_model import LogisticRegression from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score import numpy as np accuracy_lst = [] precision_lst = [] recall_lst = [] f1_lst = [] auc_lst = [] kf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42) log_reg_params = {"penalty": ['l1', 'l2'], 'C': [0.001, 0.01, 0.1, 1, 10, 100, 1000]} rand_log_reg = RandomizedSearchCV(LogisticRegression(max_iter=200), log_reg_params, n_iter=4) for train, test in kf.split(Org_X_train, Org_y_train): X_train, X_test = Org_X_train.iloc[train], Org_X_train.iloc[test] y_train, y_test = Org_y_train.iloc[train], Org_y_train.iloc[test] # 使用imblearn的Pipeline定义流程 pipeline = Pipeline([ ('oversampler', RandomOverSampler(random_state=42)), ('outlier_remover', OutlierRemover(columns=['V14', 'V12', 'V10', 'V4', 'V11', 'V2'])), ('random_search', rand_log_reg) ]) model = pipeline.fit(X_train, y_train) # 从Pipeline中获取训练后的RandomizedSearchCV实例,再取最佳估计器 best_est = model.named_steps['random_search'].best_estimator_ prediction = best_est.predict(X_test) # 计算并保存各指标 accuracy_lst.append(accuracy_score(y_test, prediction)) precision_lst.append(precision_score(y_test, prediction)) recall_lst.append(recall_score(y_test, prediction)) f1_lst.append(f1_score(y_test, prediction)) auc_lst.append(roc_auc_score(y_test, prediction)) # 输出平均指标 print("Accuracy:", np.mean(accuracy_lst)) print("Precision:", np.mean(precision_lst)) print("Recall:", np.mean(recall_lst)) print("F1 Score:", np.mean(f1_lst)) print("AUC Score:", np.mean(auc_lst))
内容的提问来源于stack exchange,提问作者Roterun
相关产品推荐
相关产品推荐

