train_test_split函数是否保持多类别数据集的类别平衡?若不能该如何处理?
Great question—this is a common concern when working with balanced multi-class datasets, so let's unpack it clearly.
1. 默认train_test_split会保持类别平衡吗?
Short answer: Not guaranteed by default.
The default behavior of train_test_split is random sampling without considering class distribution. While you might get roughly similar class ratios (especially with large datasets), there's no guarantee the split will perfectly maintain the 33% per class breakdown—smaller datasets are more prone to noticeable imbalances here.
But here's the good news: Sklearn gives you an easy fix for this.
2. 如何维持类别平衡?
If you want strict (or as close as possible) preservation of your class ratios across train/val/test sets, here are the most reliable methods:
- 使用分层拆分(Stratified Split)
The stratify parameter in train_test_split is designed exactly for this. When you pass your target labels (y) to this parameter, the split will preserve the class distribution of the original dataset.
For a 3-way split (train/val/test), you can do this in two steps:
from sklearn.model_selection import train_test_split # First split into train and temp (val + test) X_train, X_temp, y_train, y_temp = train_test_split( X, y, test_size=0.4, # Adjust based on your desired split (e.g., 60% train, 20% val, 20% test) stratify=y, random_state=42 # For reproducibility ) # Then split temp into val and test, again using stratify X_val, X_test, y_val, y_test = train_test_split( X_temp, y_temp, test_size=0.5, stratify=y_temp, random_state=42 )
This will ensure each set maintains the ~33% ratio for each class.
- 重采样(仅在训练集上操作!)
If you already have an imbalanced split (for some reason you can't use stratify), you can adjust the training set alone (never touch val/test sets to avoid data leakage):
- Oversampling: Increase the number of samples in underrepresented classes (e.g., generating synthetic samples with methods like SMOTE)
- Undersampling: Reduce the number of samples in overrepresented classes
Just remember: Resampling should only be applied to the training data to keep your validation and test sets unbiased.
- 分层K折交叉验证
If you're using cross-validation instead of a single train/val split, use StratifiedKFold instead of the standard KFold. This ensures each fold maintains the original class distribution, which is critical for reliable model evaluation.
from sklearn.model_selection import StratifiedKFold skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42) for train_idx, val_idx in skf.split(X, y): X_train, X_val = X[train_idx], X[val_idx] y_train, y_val = y[train_idx], y[val_idx] # Train and evaluate your model here
内容的提问来源于stack exchange,提问作者leila

