You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python的Scikit-learn中结合交叉验证执行SMOTE处理不平衡数据集

你提的这个问题特别关键——很多人在处理不平衡数据集时,很容易在SMOTE和交叉验证结合的环节踩数据泄露的坑!我来给你拆解正确的流程,再附上可直接运行的代码示例。

为什么你的当前代码存在局限性?

你现在的代码是先拆分单次训练/测试集,再对训练集做SMOTE,最后训练模型。这种方式只能得到一次评估结果,无法充分验证模型的泛化能力;更重要的是,如果直接把这个逻辑套用到交叉验证里(比如先对整个训练集做SMOTE再分折),会导致验证折的信息泄露到SMOTE的采样过程中,让模型评估结果过于乐观,完全不可靠。

正确的SMOTE+交叉验证流程

核心原则:SMOTE只能在交叉验证的每一轮中,仅对当前的训练折做过采样,验证折必须保持原始状态,绝对不能碰。同时因为数据集不平衡,必须用分层交叉验证来保证每折的类别分布和原始数据集一致。

方式1:用imblearn Pipeline(推荐,简洁不易出错)

imblearn提供的Pipeline会自动帮你处理“每折仅对训练数据应用SMOTE”的逻辑,避免手动操作出错:

from imblearn.pipeline import Pipeline
from imblearn.over_sampling import SMOTE
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import StratifiedKFold, cross_val_score, train_test_split

# 第一步:拆分出独立的测试集(可选但强烈推荐,用于最终无偏评估)
X_train_full, X_test, y_train_full, y_test = train_test_split(
    X, y, test_size=0.3, random_state=0, stratify=y  # stratify保证拆分后类别比例一致
)

# 构建Pipeline:先做SMOTE过采样,再训练随机森林
pipeline = Pipeline([
    ('smote', SMOTE(random_state=2)),
    ('rf', RandomForestClassifier(n_estimators=25, random_state=12))
])

# 使用分层KFold,适配不平衡数据集
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)

# 执行交叉验证
cv_scores = cross_val_score(pipeline, X_train_full, y_train_full, cv=skf, scoring='accuracy')

# 输出交叉验证结果
print(f"交叉验证准确率均值: {cv_scores.mean():.4f}")
print(f"交叉验证准确率标准差: {cv_scores.std():.4f}")

# 最后用完整训练集训练模型,在独立测试集上做最终评估
pipeline.fit(X_train_full, y_train_full)
test_accuracy = pipeline.score(X_test, y_test)
print(f"独立测试集准确率: {test_accuracy:.4f}")

方式2:手动循环实现(适合理解底层逻辑)

如果你想清晰看到每一步的操作,可以手动遍历交叉验证的每一轮:

from imblearn.over_sampling import SMOTE
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import StratifiedKFold, train_test_split
import numpy as np

# 拆分独立测试集
X_train_full, X_test, y_train_full, y_test = train_test_split(
    X, y, test_size=0.3, random_state=0, stratify=y
)

skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
cv_scores = []

for train_idx, val_idx in skf.split(X_train_full, y_train_full):
    # 获取当前折的训练/验证数据
    X_train_fold, X_val_fold = X_train_full.iloc[train_idx], X_train_full.iloc[val_idx]
    y_train_fold, y_val_fold = y_train_full.iloc[train_idx], y_train_full.iloc[val_idx]
    
    # 仅对当前训练折应用SMOTE(关键!验证折绝对不能做过采样)
    sm = SMOTE(random_state=2)
    X_train_res, y_train_res = sm.fit_resample(X_train_fold, y_train_fold.ravel())
    
    # 训练模型并在验证折上评估
    clf_rf = RandomForestClassifier(n_estimators=25, random_state=12)
    clf_rf.fit(X_train_res, y_train_res)
    val_accuracy = clf_rf.score(X_val_fold, y_val_fold)
    cv_scores.append(val_accuracy)

# 输出交叉验证结果
print(f"交叉验证准确率均值: {np.mean(cv_scores):.4f}")
print(f"交叉验证准确率标准差: {np.std(cv_scores):.4f}")

# 最终训练与测试
sm = SMOTE(random_state=2)
X_train_res_full, y_train_res_full = sm.fit_resample(X_train_full, y_train_full.ravel())
clf_rf_final = RandomForestClassifier(n_estimators=25, random_state=12)
clf_rf_final.fit(X_train_res_full, y_train_res_full)
test_accuracy = clf_rf_final.score(X_test, y_test)
print(f"独立测试集准确率: {test_accuracy:.4f}")

几个必须注意的细节

  • 绝对不能先对整个训练集做SMOTE再分折:这会导致验证折的信息泄露到采样过程中,模型评估结果完全失真。
  • 必须用StratifiedKFold:普通KFold可能会让某一折中少数类样本完全缺失,分层交叉验证能保证每折的类别比例和原始数据一致。
  • 独立测试集的作用:交叉验证是为了调参和评估模型泛化能力,而独立测试集是最终的无偏评估,避免交叉验证过程中的过拟合。

内容的提问来源于stack exchange,提问作者EmJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:12:16