You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在sklearn交叉验证中仅移除训练集内的异常值?

解决交叉验证中仅训练集移除异常值的问题

要实现仅在训练阶段移除异常值、测试集保持完整的需求,不能直接使用cross_validate的默认流程(它不会对训练集单独做预处理),推荐采用手动遍历交叉验证 fold的方式,精准控制每个训练集的异常值移除逻辑,具体实现如下:

核心思路

  1. 遍历GroupKFold生成的每个训练/测试拆分
  2. 对每个训练集单独计算异常值阈值(比如y的0.95分位数)
  3. 移除训练集中的异常样本,测试集完全保留
  4. 训练模型并计算训练/测试分数

代码实现

import numpy as np
from sklearn.model_selection import GroupKFold
from sklearn.metrics import get_scorer

# 假设你已定义好以下变量:
# reg: 你的回归模型
# X: 特征数据集(DataFrame格式)
# y_tr: 目标变量(Series格式)
# groups: 分组变量(对应GroupKFold的分组)
# scoring: 评估指标(比如"r2")

# 初始化分组交叉验证器
gkf = GroupKFold(n_splits=3)
# 获取评分器
scorer = get_scorer(scoring)

# 存储训练和测试分数
train_scores = []
test_scores = []

# 遍历每个fold
for train_idx, test_idx in gkf.split(X, y_tr, groups=groups):
    # 拆分训练集和测试集
    X_train, X_test = X.iloc[train_idx], X.iloc[test_idx]
    y_train, y_test = y_tr.iloc[train_idx], y_tr.iloc[test_idx]
    
    # 计算训练集y的0.95分位数(根据需求调整分位数)
    threshold = np.quantile(y_train, 0.95)
    
    # 筛选训练集:移除y高于阈值的异常样本(如果要移除低于阈值的,改为y_train >= threshold)
    mask = y_train <= threshold
    X_train_clean = X_train[mask]
    y_train_clean = y_train[mask]
    
    # 训练模型
    reg.fit(X_train_clean, y_train_clean)
    
    # 计算分数
    train_score = scorer(reg, X_train_clean, y_train_clean)
    test_score = scorer(reg, X_test, y_test)
    
    train_scores.append(train_score)
    test_scores.append(test_score)

# 整理成和cross_validate一致的输出格式
cv_scores = {
    'train_score': train_scores,
    'test_score': test_scores
}

针对你的示例数据的说明

以训练集包含id a和b为例:

  • 训练集y值为[100, 150, 130, 1000, 90],排序后为[90, 100, 130, 150, 1000]
  • 0.95分位数为957.5,mask = y_train <= threshold会筛选掉y=1000的样本(即date 2 id b的异常样本)
  • 测试集(id c)的样本完全保留,符合你的需求

替代方案说明

如果你更倾向于用sklearn的Pipeline封装流程,可以自定义一个仅在训练阶段生效的异常值移除器(需要同时处理X和y),但手动遍历的方式更直观、易调试,更适配你的场景。

内容的提问来源于stack exchange,提问作者Fernando Quintino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 17:25:50