HistGradientBoostingClassifier训练报错:y中存在仅含1个样本的类别
问题描述
使用新冠病例与死亡关联数据集构建HistGradientBoostingClassifier模型时,调用fit方法触发ValueError,提示y中最不常见类别仅含1个样本,无法满足分层抽样至少2个样本的要求,但未发现X、y拆分逻辑存在问题,也不完全理解该错误。
代码示例
import numpy as np #from sklearn.tree import DecisionTreeClassifier from sklearn.ensemble import HistGradientBoostingClassifier from sklearn import preprocessing #import csv X_test = pd.read_csv("test.csv") y_output = pd.read_csv("sample_submission.csv") data_train = pd.read_csv("train.csv") X_train = data_train.drop(columns=["Next Week's Deaths"]) y_train = data_train["Next Week's Deaths"] #prepare for fit (transform Location strings into classes) Location = data_train["Location"] le = preprocessing.LabelEncoder() le.fit(Location) LocationToInt = le.transform(Location) LocationDict = dict(zip(Location, LocationToInt)) X_train["Location"] = X_train["Location"].replace(LocationDict) #train and run model = HistGradientBoostingClassifier(max_bins=255, max_iter=100) model.fit(X_train, y_train)
报错栈
Input In [89], in <cell line: 29>() 27 #train and run 28 model = HistGradientBoostingClassifier(max_bins=255, max_iter=100) ---> 29 model.fit(X_train, y_train) File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\ensemble\_hist_gradient_boosting\gradient_boosting.py:348, in BaseHistGradientBoosting.fit(self, X, y, sample_weight) 343 # Save the state of the RNG for the training and validation split. 344 # This is needed in order to have the same split when using 345 # warm starting. 347 if sample_weight is None: --> 348 X_train, X_val, y_train, y_val = train_test_split( 349 X, 350 y, 351 test_size=self.validation_fraction, 352 stratify=stratify, 353 random_state=self._random_seed, 354 ) 355 sample_weight_train = sample_weight_val = None 356 else: 357 # TODO: incorporate sample_weight in sampling here, as well as 358 # stratify File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\model_selection\_split.py:2454, in train_test_split(test_size, train_size, random_state, shuffle, stratify, *arrays) 2450 CVClass = ShuffleSplit 2452 cv = CVClass(test_size=n_test, train_size=n_train, random_state=random_state) -> 2454 train, test = next(cv.split(X=arrays[0], y=stratify)) 2456 return list( 2457 chain.from_iterable( 2458 (_safe_indexing(a, train), _safe_indexing(a, test)) for a in arrays 2459 ) 2460 ) File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\model_selection\_split.py:1613, in BaseShuffleSplit.split(self, X, y, groups) 1583 """Generate indices to split data into training and test set. 1584 1585 Parameters (...) 1610 to an integer. 1611 """ 1612 X, y, groups = indexable(X, y, groups) -> 1613 for train, test in self._iter_indices(X, y, groups): 1614 yield train, test File ~\AppData\Local\Programs\Python\Python310\lib\site-packages\sklearn\model_selection\_split.py:1953, in StratifiedShuffleSplit._iter_indices(self, X, y, groups) 1951 class_counts = np.bincount(y_indices) 1952 if np.min(class_counts) < 2: -> 1953 raise ValueError( 1954 "The least populated class in y has only 1" 1955 " member, which is too few. The minimum" 1956 " number of groups for any class cannot" 1957 " be less than 2." 1958 ) 1960 if n_train < n_classes: 1961 raise ValueError( 1962 "The train_size = %d should be greater or " 1963 "equal to the number of classes = %d" % (n_train, n_classes) 1964 ) ValueError: The least populated class in y has only 1 member, which is too few. The minimum number of groups for any class cannot be less than 2.
问题原因及解决办法
原因
HistGradientBoostingClassifier默认会使用分层抽样划分训练集和验证集(validation_fraction默认值0.1),而你的标签y_train中存在某个类别只有1个样本,分层抽样要求每个类别至少有2个样本(才能分到训练和验证集各至少1个),因此触发错误。
解决办法
有三种可行方案:
- 关闭分层抽样:初始化模型时设置
stratify=False,让模型使用随机抽样划分验证集,不再强制分层:
model = HistGradientBoostingClassifier(max_bins=255, max_iter=100, stratify=False)
- 移除样本量不足的类别:先检查
y_train中各类别的样本数,删除只有1个样本的类别对应的行:
# 统计每个类别的样本数 class_counts = y_train.value_counts() # 筛选出样本数>=2的类别 valid_classes = class_counts[class_counts >= 2].index # 过滤训练数据 X_train = X_train[y_train.isin(valid_classes)] y_train = y_train[y_train.isin(valid_classes)]
- 关闭内置验证集:设置
validation_fraction=0,模型不再划分验证集,直接用全部数据训练:
model = HistGradientBoostingClassifier(max_bins=255, max_iter=100, validation_fraction=0)
内容的提问来源于stack exchange,提问作者Felix Kniest
相关产品推荐
相关产品推荐

