如何在sklearn交叉验证中仅移除训练集内的异常值?
解决交叉验证中仅训练集移除异常值的问题
要实现仅在训练阶段移除异常值、测试集保持完整的需求,不能直接使用cross_validate的默认流程(它不会对训练集单独做预处理),推荐采用手动遍历交叉验证 fold的方式,精准控制每个训练集的异常值移除逻辑,具体实现如下:
核心思路
- 遍历
GroupKFold生成的每个训练/测试拆分 - 对每个训练集单独计算异常值阈值(比如y的0.95分位数)
- 移除训练集中的异常样本,测试集完全保留
- 训练模型并计算训练/测试分数
代码实现
import numpy as np from sklearn.model_selection import GroupKFold from sklearn.metrics import get_scorer # 假设你已定义好以下变量: # reg: 你的回归模型 # X: 特征数据集(DataFrame格式) # y_tr: 目标变量(Series格式) # groups: 分组变量(对应GroupKFold的分组) # scoring: 评估指标(比如"r2") # 初始化分组交叉验证器 gkf = GroupKFold(n_splits=3) # 获取评分器 scorer = get_scorer(scoring) # 存储训练和测试分数 train_scores = [] test_scores = [] # 遍历每个fold for train_idx, test_idx in gkf.split(X, y_tr, groups=groups): # 拆分训练集和测试集 X_train, X_test = X.iloc[train_idx], X.iloc[test_idx] y_train, y_test = y_tr.iloc[train_idx], y_tr.iloc[test_idx] # 计算训练集y的0.95分位数(根据需求调整分位数) threshold = np.quantile(y_train, 0.95) # 筛选训练集:移除y高于阈值的异常样本(如果要移除低于阈值的,改为y_train >= threshold) mask = y_train <= threshold X_train_clean = X_train[mask] y_train_clean = y_train[mask] # 训练模型 reg.fit(X_train_clean, y_train_clean) # 计算分数 train_score = scorer(reg, X_train_clean, y_train_clean) test_score = scorer(reg, X_test, y_test) train_scores.append(train_score) test_scores.append(test_score) # 整理成和cross_validate一致的输出格式 cv_scores = { 'train_score': train_scores, 'test_score': test_scores }
针对你的示例数据的说明
以训练集包含id a和b为例:
- 训练集y值为
[100, 150, 130, 1000, 90],排序后为[90, 100, 130, 150, 1000] - 0.95分位数为
957.5,mask = y_train <= threshold会筛选掉y=1000的样本(即date 2 id b的异常样本) - 测试集(id c)的样本完全保留,符合你的需求
替代方案说明
如果你更倾向于用sklearn的Pipeline封装流程,可以自定义一个仅在训练阶段生效的异常值移除器(需要同时处理X和y),但手动遍历的方式更直观、易调试,更适配你的场景。
内容的提问来源于stack exchange,提问作者Fernando Quintino
相关产品推荐
相关产品推荐

