如何在Python的Scikit-learn中结合交叉验证执行SMOTE处理不平衡数据集
你提的这个问题特别关键——很多人在处理不平衡数据集时,很容易在SMOTE和交叉验证结合的环节踩数据泄露的坑!我来给你拆解正确的流程,再附上可直接运行的代码示例。
为什么你的当前代码存在局限性?
你现在的代码是先拆分单次训练/测试集,再对训练集做SMOTE,最后训练模型。这种方式只能得到一次评估结果,无法充分验证模型的泛化能力;更重要的是,如果直接把这个逻辑套用到交叉验证里(比如先对整个训练集做SMOTE再分折),会导致验证折的信息泄露到SMOTE的采样过程中,让模型评估结果过于乐观,完全不可靠。
正确的SMOTE+交叉验证流程
核心原则:SMOTE只能在交叉验证的每一轮中,仅对当前的训练折做过采样,验证折必须保持原始状态,绝对不能碰。同时因为数据集不平衡,必须用分层交叉验证来保证每折的类别分布和原始数据集一致。
方式1:用imblearn Pipeline(推荐,简洁不易出错)
imblearn提供的Pipeline会自动帮你处理“每折仅对训练数据应用SMOTE”的逻辑,避免手动操作出错:
from imblearn.pipeline import Pipeline from imblearn.over_sampling import SMOTE from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import StratifiedKFold, cross_val_score, train_test_split # 第一步:拆分出独立的测试集(可选但强烈推荐,用于最终无偏评估) X_train_full, X_test, y_train_full, y_test = train_test_split( X, y, test_size=0.3, random_state=0, stratify=y # stratify保证拆分后类别比例一致 ) # 构建Pipeline:先做SMOTE过采样,再训练随机森林 pipeline = Pipeline([ ('smote', SMOTE(random_state=2)), ('rf', RandomForestClassifier(n_estimators=25, random_state=12)) ]) # 使用分层KFold,适配不平衡数据集 skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=0) # 执行交叉验证 cv_scores = cross_val_score(pipeline, X_train_full, y_train_full, cv=skf, scoring='accuracy') # 输出交叉验证结果 print(f"交叉验证准确率均值: {cv_scores.mean():.4f}") print(f"交叉验证准确率标准差: {cv_scores.std():.4f}") # 最后用完整训练集训练模型,在独立测试集上做最终评估 pipeline.fit(X_train_full, y_train_full) test_accuracy = pipeline.score(X_test, y_test) print(f"独立测试集准确率: {test_accuracy:.4f}")
方式2:手动循环实现(适合理解底层逻辑)
如果你想清晰看到每一步的操作,可以手动遍历交叉验证的每一轮:
from imblearn.over_sampling import SMOTE from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import StratifiedKFold, train_test_split import numpy as np # 拆分独立测试集 X_train_full, X_test, y_train_full, y_test = train_test_split( X, y, test_size=0.3, random_state=0, stratify=y ) skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=0) cv_scores = [] for train_idx, val_idx in skf.split(X_train_full, y_train_full): # 获取当前折的训练/验证数据 X_train_fold, X_val_fold = X_train_full.iloc[train_idx], X_train_full.iloc[val_idx] y_train_fold, y_val_fold = y_train_full.iloc[train_idx], y_train_full.iloc[val_idx] # 仅对当前训练折应用SMOTE(关键!验证折绝对不能做过采样) sm = SMOTE(random_state=2) X_train_res, y_train_res = sm.fit_resample(X_train_fold, y_train_fold.ravel()) # 训练模型并在验证折上评估 clf_rf = RandomForestClassifier(n_estimators=25, random_state=12) clf_rf.fit(X_train_res, y_train_res) val_accuracy = clf_rf.score(X_val_fold, y_val_fold) cv_scores.append(val_accuracy) # 输出交叉验证结果 print(f"交叉验证准确率均值: {np.mean(cv_scores):.4f}") print(f"交叉验证准确率标准差: {np.std(cv_scores):.4f}") # 最终训练与测试 sm = SMOTE(random_state=2) X_train_res_full, y_train_res_full = sm.fit_resample(X_train_full, y_train_full.ravel()) clf_rf_final = RandomForestClassifier(n_estimators=25, random_state=12) clf_rf_final.fit(X_train_res_full, y_train_res_full) test_accuracy = clf_rf_final.score(X_test, y_test) print(f"独立测试集准确率: {test_accuracy:.4f}")
几个必须注意的细节
- 绝对不能先对整个训练集做SMOTE再分折:这会导致验证折的信息泄露到采样过程中,模型评估结果完全失真。
- 必须用
StratifiedKFold:普通KFold可能会让某一折中少数类样本完全缺失,分层交叉验证能保证每折的类别比例和原始数据一致。 - 独立测试集的作用:交叉验证是为了调参和评估模型泛化能力,而独立测试集是最终的无偏评估,避免交叉验证过程中的过拟合。
内容的提问来源于stack exchange,提问作者EmJ
相关产品推荐
相关产品推荐

