如何将预计算的欠采样交叉验证折传入sklearn的HalvingRandomSearchCV
预计算欠采样交叉验证折优化GBM模型调参效率
问题描述
需求是使用imblearn.RandomUnderSampler生成3个欠采样交叉验证折,基于这些折对XGBoost、CatBoost、GradientBoostingClassifier等GBM模型进行超参数优化。现有代码在每次超参数搜索时都会重复执行欠采样操作,导致效率低下,希望将预计算好的欠采样折直接传入HalvingRandomSearchCV。
可行实现方案
HalvingRandomSearchCV的cv参数支持传入自定义的交叉验证迭代器(由(训练索引, 测试索引)元组组成的列表),我们可以预先对每个训练折完成欠采样,再将处理后的索引对传入cv,从而避免重复欠采样。
核心思路
- 先用普通
KFold拆分原始训练集为3个基础折,每个折包含原始训练索引和原始测试索引 - 对每个基础折的训练索引对应的子数据集执行欠采样,得到欠采样后的训练索引
- 将每个折替换为
(欠采样训练索引, 原始测试索引),组成预计算折列表传入HalvingRandomSearchCV的cv参数 - 移除原Pipeline中的欠采样步骤,仅保留必要的特征缩放和模型组件
代码实现
1. 生成预计算欠采样交叉验证折
from sklearn.model_selection import KFold from imblearn.under_sampling import RandomUnderSampler def generate_undersampled_cv_folds(X_train, y_train, n_splits=3, random_state=10): # 初始化KFold拆分原始数据集 kf = KFold(n_splits=n_splits, shuffle=True, random_state=random_state) undersampled_folds = [] for train_idx, test_idx in kf.split(X_train): # 提取当前训练折的子数据集,兼容DataFrame和数组格式 X_fold_train = X_train.iloc[train_idx] if hasattr(X_train, 'iloc') else X_train[train_idx] y_fold_train = y_train.iloc[train_idx] if hasattr(y_train, 'iloc') else y_train[train_idx] # 执行欠采样,获取欠采样后的样本索引 rus = RandomUnderSampler(random_state=random_state) resampled_idx, _ = rus.fit_resample( X_fold_train.index.values.reshape(-1, 1) if hasattr(X_fold_train, 'index') else train_idx.reshape(-1, 1), y_fold_train ) resampled_train_idx = resampled_idx.flatten() # 保存欠采样训练索引+原始测试索引的折 undersampled_folds.append((resampled_train_idx, test_idx)) return undersampled_folds
2. 修改后的模型训练与超参数搜索函数
from sklearn.pipeline import Pipeline from sklearn.preprocessing import MinMaxScaler from sklearn.model_selection import HalvingRandomSearchCV def train_model_with_precomputed_folds(estimator, scale, params, X_train, y_train, cv_folds): # 构建仅包含特征缩放(可选)和模型的Pipeline if scale is True: pipe = Pipeline([ ("scaler", MinMaxScaler()), ("model", estimator), ]) else: pipe = Pipeline([ ("model", estimator), ]) search = HalvingRandomSearchCV( estimator=pipe, param_distributions=params, n_candidates="exhaust", factor=3, resource='model__n_estimators', max_resources=500, min_resources=10, scoring='roc_auc', cv=cv_folds, # 使用预计算的欠采样折 random_state=10, refit=True, n_jobs=-1, ) search.fit(X_train, y_train) return search
3. 使用示例(以XGBoost为例)
from xgboost import XGBClassifier # 生成预计算欠采样折 cv_folds = generate_undersampled_cv_folds(X_train, y_train) # 定义超参数搜索空间 params = { "model__max_depth": [3, 5, 7], "model__learning_rate": [0.01, 0.1, 0.2], "model__subsample": [0.8, 1.0], # 可添加更多GBM模型的超参数 } # 执行超参数搜索 search = train_model_with_precomputed_folds( estimator=XGBClassifier(), scale=True, params=params, X_train=X_train, y_train=y_train, cv_folds=cv_folds )
关键注意事项
- 测试折不做欠采样:必须保留原始测试折的数据分布,这样模型评估结果才能真实反映其在不平衡真实数据上的性能
- 索引兼容性:代码同时支持Pandas DataFrame和Numpy数组输入,通过判断
iloc和index属性适配不同数据类型 - 可重复性:固定
random_state保证欠采样和交叉验证折的可复现性
内容的提问来源于stack exchange,提问作者Sole Galli
相关产品推荐
相关产品推荐

