You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将预计算的欠采样交叉验证折传入sklearn的HalvingRandomSearchCV

预计算欠采样交叉验证折优化GBM模型调参效率

问题描述

需求是使用imblearn.RandomUnderSampler生成3个欠采样交叉验证折,基于这些折对XGBoost、CatBoost、GradientBoostingClassifier等GBM模型进行超参数优化。现有代码在每次超参数搜索时都会重复执行欠采样操作,导致效率低下,希望将预计算好的欠采样折直接传入HalvingRandomSearchCV。

可行实现方案

HalvingRandomSearchCV的cv参数支持传入自定义的交叉验证迭代器(由(训练索引, 测试索引)元组组成的列表),我们可以预先对每个训练折完成欠采样,再将处理后的索引对传入cv,从而避免重复欠采样。

核心思路

  1. 先用普通KFold拆分原始训练集为3个基础折,每个折包含原始训练索引和原始测试索引
  2. 对每个基础折的训练索引对应的子数据集执行欠采样,得到欠采样后的训练索引
  3. 将每个折替换为(欠采样训练索引, 原始测试索引),组成预计算折列表传入HalvingRandomSearchCV的cv参数
  4. 移除原Pipeline中的欠采样步骤,仅保留必要的特征缩放和模型组件

代码实现

1. 生成预计算欠采样交叉验证折

from sklearn.model_selection import KFold
from imblearn.under_sampling import RandomUnderSampler

def generate_undersampled_cv_folds(X_train, y_train, n_splits=3, random_state=10):
    # 初始化KFold拆分原始数据集
    kf = KFold(n_splits=n_splits, shuffle=True, random_state=random_state)
    undersampled_folds = []
    
    for train_idx, test_idx in kf.split(X_train):
        # 提取当前训练折的子数据集,兼容DataFrame和数组格式
        X_fold_train = X_train.iloc[train_idx] if hasattr(X_train, 'iloc') else X_train[train_idx]
        y_fold_train = y_train.iloc[train_idx] if hasattr(y_train, 'iloc') else y_train[train_idx]
        
        # 执行欠采样,获取欠采样后的样本索引
        rus = RandomUnderSampler(random_state=random_state)
        resampled_idx, _ = rus.fit_resample(
            X_fold_train.index.values.reshape(-1, 1) if hasattr(X_fold_train, 'index') else train_idx.reshape(-1, 1),
            y_fold_train
        )
        resampled_train_idx = resampled_idx.flatten()
        
        # 保存欠采样训练索引+原始测试索引的折
        undersampled_folds.append((resampled_train_idx, test_idx))
    
    return undersampled_folds

2. 修改后的模型训练与超参数搜索函数

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import MinMaxScaler
from sklearn.model_selection import HalvingRandomSearchCV

def train_model_with_precomputed_folds(estimator, scale, params, X_train, y_train, cv_folds):
    # 构建仅包含特征缩放(可选)和模型的Pipeline
    if scale is True:
        pipe = Pipeline([
            ("scaler", MinMaxScaler()),
            ("model", estimator),
        ])
    else:
        pipe = Pipeline([
            ("model", estimator),
        ])

    search = HalvingRandomSearchCV(
        estimator=pipe,
        param_distributions=params,
        n_candidates="exhaust",
        factor=3,
        resource='model__n_estimators',
        max_resources=500,
        min_resources=10,
        scoring='roc_auc',
        cv=cv_folds,  # 使用预计算的欠采样折
        random_state=10,
        refit=True,
        n_jobs=-1,
    )

    search.fit(X_train, y_train)
    return search

3. 使用示例(以XGBoost为例)

from xgboost import XGBClassifier

# 生成预计算欠采样折
cv_folds = generate_undersampled_cv_folds(X_train, y_train)

# 定义超参数搜索空间
params = {
    "model__max_depth": [3, 5, 7],
    "model__learning_rate": [0.01, 0.1, 0.2],
    "model__subsample": [0.8, 1.0],
    # 可添加更多GBM模型的超参数
}

# 执行超参数搜索
search = train_model_with_precomputed_folds(
    estimator=XGBClassifier(),
    scale=True,
    params=params,
    X_train=X_train,
    y_train=y_train,
    cv_folds=cv_folds
)

关键注意事项

  • 测试折不做欠采样:必须保留原始测试折的数据分布,这样模型评估结果才能真实反映其在不平衡真实数据上的性能
  • 索引兼容性:代码同时支持Pandas DataFrame和Numpy数组输入,通过判断iloc和index属性适配不同数据类型
  • 可重复性:固定random_state保证欠采样和交叉验证折的可复现性

内容的提问来源于stack exchange,提问作者Sole Galli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 11:27:35