You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决Python ValueError:训练测试集拆分时最小类别样本数不足问题

解决方法:拆分训练测试集时遇ValueError(类别样本量过少)

错误原因

你用train_test_split指定stratify=y做分层拆分时,目标变量y里存在仅含1个样本的类别——分层拆分需要每个类别至少有2个样本,才能分配到训练集和测试集里,所以触发报错。

具体解决方案

  • 先清理/合并稀有类别
    先查看类别分布找到问题类别,再处理:

    # 统计每个类别的样本数量
    print(y.value_counts())
    # 筛选出样本数≤1的类别
    rare_classes = y.value_counts()[y.value_counts() <= 1].index
    

    处理方式二选一:

    1. 删除稀有类别样本:
      # 过滤掉稀有类别的数据行
      filtered_data = df[~df['target_col'].isin(rare_classes)]
      X = filtered_data.drop('target_col', axis=1)
      y = filtered_data['target_col']
      
    2. 合并到相似类别(根据业务逻辑调整):
      # 把稀有类别统一替换为"other"
      y = y.replace(rare_classes, 'other')
      
  • 取消分层拆分(适合对类别分布要求不高的场景)
    如果不需要严格保持训练/测试集的类别分布一致,直接去掉stratify=y参数:

    from sklearn.model_selection import train_test_split
    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
    
  • 调整分层抽样策略(小样本场景)
    若必须分层,需确保所有类别样本数≥拆分折数,再用分层逻辑:

    from sklearn.model_selection import StratifiedKFold
    # 最低要求n_splits设为2,确保每个类别至少能分到两个集合
    skf = StratifiedKFold(n_splits=2, shuffle=True, random_state=42)
    for train_idx, test_idx in skf.split(X, y):
        X_train, X_test = X.iloc[train_idx], X.iloc[test_idx]
        y_train, y_test = y.iloc[train_idx], y.iloc[test_idx]
    

内容的提问来源于stack exchange,提问作者Niraja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 09:55:18