解决Python ValueError:训练测试集拆分时最小类别样本数不足问题
解决方法:拆分训练测试集时遇ValueError(类别样本量过少)
错误原因
你用train_test_split指定stratify=y做分层拆分时,目标变量y里存在仅含1个样本的类别——分层拆分需要每个类别至少有2个样本,才能分配到训练集和测试集里,所以触发报错。
具体解决方案
先清理/合并稀有类别
先查看类别分布找到问题类别,再处理:# 统计每个类别的样本数量 print(y.value_counts()) # 筛选出样本数≤1的类别 rare_classes = y.value_counts()[y.value_counts() <= 1].index处理方式二选一:
- 删除稀有类别样本:
# 过滤掉稀有类别的数据行 filtered_data = df[~df['target_col'].isin(rare_classes)] X = filtered_data.drop('target_col', axis=1) y = filtered_data['target_col'] - 合并到相似类别(根据业务逻辑调整):
# 把稀有类别统一替换为"other" y = y.replace(rare_classes, 'other')
- 删除稀有类别样本:
取消分层拆分(适合对类别分布要求不高的场景)
如果不需要严格保持训练/测试集的类别分布一致,直接去掉stratify=y参数:from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)调整分层抽样策略(小样本场景)
若必须分层,需确保所有类别样本数≥拆分折数,再用分层逻辑:from sklearn.model_selection import StratifiedKFold # 最低要求n_splits设为2,确保每个类别至少能分到两个集合 skf = StratifiedKFold(n_splits=2, shuffle=True, random_state=42) for train_idx, test_idx in skf.split(X, y): X_train, X_test = X.iloc[train_idx], X.iloc[test_idx] y_train, y_test = y.iloc[train_idx], y.iloc[test_idx]
内容的提问来源于stack exchange,提问作者Niraja
相关产品推荐
相关产品推荐

