You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

train_test_split函数是否保持多类别数据集的类别平衡?若不能该如何处理?

关于train_test_split类别平衡的问题解答

Great question—this is a common concern when working with balanced multi-class datasets, so let's unpack it clearly.

1. 默认train_test_split会保持类别平衡吗?

Short answer: Not guaranteed by default.

The default behavior of train_test_split is random sampling without considering class distribution. While you might get roughly similar class ratios (especially with large datasets), there's no guarantee the split will perfectly maintain the 33% per class breakdown—smaller datasets are more prone to noticeable imbalances here.

But here's the good news: Sklearn gives you an easy fix for this.

2. 如何维持类别平衡?

If you want strict (or as close as possible) preservation of your class ratios across train/val/test sets, here are the most reliable methods:

- 使用分层拆分(Stratified Split)

The stratify parameter in train_test_split is designed exactly for this. When you pass your target labels (y) to this parameter, the split will preserve the class distribution of the original dataset.

For a 3-way split (train/val/test), you can do this in two steps:

from sklearn.model_selection import train_test_split

# First split into train and temp (val + test)
X_train, X_temp, y_train, y_temp = train_test_split(
    X, y, 
    test_size=0.4,  # Adjust based on your desired split (e.g., 60% train, 20% val, 20% test)
    stratify=y,
    random_state=42  # For reproducibility
)

# Then split temp into val and test, again using stratify
X_val, X_test, y_val, y_test = train_test_split(
    X_temp, y_temp, 
    test_size=0.5,
    stratify=y_temp,
    random_state=42
)

This will ensure each set maintains the ~33% ratio for each class.

- 重采样(仅在训练集上操作!)

If you already have an imbalanced split (for some reason you can't use stratify), you can adjust the training set alone (never touch val/test sets to avoid data leakage):

  • Oversampling: Increase the number of samples in underrepresented classes (e.g., generating synthetic samples with methods like SMOTE)
  • Undersampling: Reduce the number of samples in overrepresented classes

Just remember: Resampling should only be applied to the training data to keep your validation and test sets unbiased.

- 分层K折交叉验证

If you're using cross-validation instead of a single train/val split, use StratifiedKFold instead of the standard KFold. This ensures each fold maintains the original class distribution, which is critical for reliable model evaluation.

from sklearn.model_selection import StratifiedKFold

skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
for train_idx, val_idx in skf.split(X, y):
    X_train, X_val = X[train_idx], X[val_idx]
    y_train, y_val = y[train_idx], y[val_idx]
    # Train and evaluate your model here

内容的提问来源于stack exchange,提问作者leila

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:52:36