You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HistGradientBoostingClassifier训练报错:y中存在仅含1个样本的类别

问题描述

使用新冠病例与死亡关联数据集构建HistGradientBoostingClassifier模型时,调用fit方法触发ValueError,提示y中最不常见类别仅含1个样本,无法满足分层抽样至少2个样本的要求,但未发现X、y拆分逻辑存在问题,也不完全理解该错误。

代码示例

import numpy as np
#from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn import preprocessing


#import csv
X_test = pd.read_csv("test.csv")
y_output = pd.read_csv("sample_submission.csv")

data_train = pd.read_csv("train.csv")
X_train = data_train.drop(columns=["Next Week's Deaths"])
y_train = data_train["Next Week's Deaths"]

#prepare for fit (transform Location strings into classes)
Location = data_train["Location"]
le = preprocessing.LabelEncoder()
le.fit(Location)

LocationToInt = le.transform(Location)
LocationDict = dict(zip(Location, LocationToInt))

X_train["Location"] = X_train["Location"].replace(LocationDict)


#train and run
model = HistGradientBoostingClassifier(max_bins=255, max_iter=100)
model.fit(X_train, y_train)

报错栈

Input In [89], in <cell line: 29>()
     27 #train and run
     28 model = HistGradientBoostingClassifier(max_bins=255, max_iter=100)
---&gt; 29 model.fit(X_train, y_train)

File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\ensemble\_hist_gradient_boosting\gradient_boosting.py:348, in BaseHistGradientBoosting.fit(self, X, y, sample_weight)
    343 # Save the state of the RNG for the training and validation split.
    344 # This is needed in order to have the same split when using
    345 # warm starting.
    347 if sample_weight is None:
--&gt; 348     X_train, X_val, y_train, y_val = train_test_split(
    349         X,
    350         y,
    351         test_size=self.validation_fraction,
    352         stratify=stratify,
    353         random_state=self._random_seed,
    354     )
    355     sample_weight_train = sample_weight_val = None
    356 else:
    357     # TODO: incorporate sample_weight in sampling here, as well as
    358     # stratify

File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\model_selection\_split.py:2454, in train_test_split(test_size, train_size, random_state, shuffle, stratify, *arrays)
   2450         CVClass = ShuffleSplit
   2452     cv = CVClass(test_size=n_test, train_size=n_train, random_state=random_state)
-&gt; 2454     train, test = next(cv.split(X=arrays[0], y=stratify))
   2456 return list(
   2457     chain.from_iterable(
   2458         (_safe_indexing(a, train), _safe_indexing(a, test)) for a in arrays
   2459     )
   2460 )

File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\model_selection\_split.py:1613, in BaseShuffleSplit.split(self, X, y, groups)
   1583 """Generate indices to split data into training and test set.
   1584 
   1585 Parameters
   (...)
   1610 to an integer.
   1611 """
   1612 X, y, groups = indexable(X, y, groups)
-&gt; 1613 for train, test in self._iter_indices(X, y, groups):
   1614     yield train, test

File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\model_selection\_split.py:1953, in StratifiedShuffleSplit._iter_indices(self, X, y, groups)
   1951 class_counts = np.bincount(y_indices)
   1952 if np.min(class_counts) &lt; 2:
-&gt; 1953     raise ValueError(
   1954         "The least populated class in y has only 1"
   1955         " member, which is too few. The minimum"
   1956         " number of groups for any class cannot"
   1957         " be less than 2."
   1958     )
   1960 if n_train &lt; n_classes:
   1961     raise ValueError(
   1962         "The train_size = %d should be greater or "
   1963         "equal to the number of classes = %d" % (n_train, n_classes)
   1964     )

ValueError: The least populated class in y has only 1 member, which is too few. The minimum number of groups for any class cannot be less than 2.
问题原因及解决办法

原因

HistGradientBoostingClassifier默认会使用分层抽样划分训练集和验证集(validation_fraction默认值0.1),而你的标签y_train中存在某个类别只有1个样本,分层抽样要求每个类别至少有2个样本(才能分到训练和验证集各至少1个),因此触发错误。

解决办法

有三种可行方案:

  • 关闭分层抽样:初始化模型时设置stratify=False,让模型使用随机抽样划分验证集,不再强制分层:
model = HistGradientBoostingClassifier(max_bins=255, max_iter=100, stratify=False)
  • 移除样本量不足的类别:先检查y_train中各类别的样本数,删除只有1个样本的类别对应的行:
# 统计每个类别的样本数
class_counts = y_train.value_counts()
# 筛选出样本数>=2的类别
valid_classes = class_counts[class_counts >= 2].index
# 过滤训练数据
X_train = X_train[y_train.isin(valid_classes)]
y_train = y_train[y_train.isin(valid_classes)]
  • 关闭内置验证集:设置validation_fraction=0,模型不再划分验证集,直接用全部数据训练:
model = HistGradientBoostingClassifier(max_bins=255, max_iter=100, validation_fraction=0)

内容的提问来源于stack exchange,提问作者Felix Kniest

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 18:27:29