You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用GridSearchCV fit遇ValueError:n_splits=5大于各类成员数

问题原因分析:GridSearchCV fit时的ValueError错误

问题触发场景

调用GridSearchCV的fit方法时出现如下错误:

ValueError: n_splits=5 cannot be greater than the number of members in each class.

相关输入信息

训练特征集信息(X_train.info()输出)

<class 'pandas.core.frame.DataFrame'>
Int64Index: 6000 entries, 9761 to 7270
Data columns (total 4 columns):
 #   Column  Non-Null Count  Dtype
---  ------  --------------  -----
 0   a       6000 non-null   int64
 1   b       6000 non-null   int64
 2   c       6000 non-null   int64
 3   d       6000 non-null   int64
dtypes: int64(4)
memory usage: 234.4 KB

训练标签集信息(y_train.info()输出)

<class 'pandas.core.series.Series'>
Int64Index: 6000 entries, 9761 to 7270
Series name: result
Non-Null Count  Dtype
--------------  -----
6000 non-null   int64
dtypes: int64(1)
memory usage: 93.8 KB

GridSearchCV调用代码

parameters = {
    "max_depth": [1, 2, 3],
}

cv = GridSearchCV(
    DecisionTreeClassifier(), 
    parameters, 
    cv=5,
    verbose=1,
)

cv.fit(X_train, y_train)

错误核心原因

训练集标签(y_train)存在严重类别不平衡,至少有一个类别的样本总数小于5。

GridSearchCV默认采用分层K折交叉验证(StratifiedKFold),这种策略要求每个折中的类别分布与整体数据集一致。当设置cv=5时,每个类别需要能被分到5个折里,每个折至少包含该类别1个样本。若某个类别的总样本数不足5,无法满足这个分配要求,就会触发该错误。

你可以执行以下代码查看各类别样本数量:

print(y_train.value_counts())

可选解决方向

  • 降低交叉验证折数:将cv设置为小于等于最小类别样本数的数值(比如2、3)
  • 样本重采样:对少数类别进行过采样,或对多数类别进行欠采样,平衡类别分布
  • 更换交叉验证策略:使用非分层的KFold,但这种方法可能在不平衡数据集上导致验证结果偏差

内容的提问来源于stack exchange,提问作者e42d3

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 16:25:16