You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SKLearn的CategoricalNB时遭遇IndexError问题求助

IndexError with CategoricalNB.predict() when using single features in K-Fold cross-validation

问题背景

使用Scikit-learn的CategoricalNB实现朴素贝叶斯分类,进行前向选择的单特征训练时,第一个特征可正常完成10折交叉验证,但遍历第二个特征时,clf.predict()触发如下错误:

IndexError: index 29 is out of bounds for axis 1 with size 29

已尝试仅传入单个特征列、设置min_categories为特征唯一类别数,问题仍未解决。

错误原因

  1. 核心问题:原始分类特征未做整数编码
    CategoricalNB要求输入特征必须是从0开始的连续整数类型的类别索引,不能直接传入原始分类值(如字符串、非连续整数ID)。当直接传入原始特征值时,模型会将该值当作类别索引访问feature_log_prob_数组,若特征值大于等于模型推断的类别数(比如训练集该特征的最大索引为28,但测试集出现29),就会触发索引越界。
  2. 次要问题:KFold拆分对象不匹配
    代码中用全特征集X执行kf.split(X),虽然不会直接导致索引错误,但逻辑上应使用当前的单特征集singleX,保持数据拆分的一致性。

修复方案

1. 对分类特征进行整数编码

使用OrdinalEncoder(专门用于特征列的编码工具)将原始类别转换为从0开始的连续整数,且基于整个数据集拟合编码器,确保覆盖所有可能的类别,避免测试集出现训练集未见过的类别。

2. 调整KFold拆分对象

用当前单特征集singleX替代全特征集X执行拆分,逻辑更严谨。

3. 合理设置min_categories(可选)

若需兼容训练集可能缺失的类别,可基于整个数据集的类别数设置min_categories,但前提是编码已覆盖所有类别。

修改后的完整代码

import pandas as pd
from sklearn.naive_bayes import CategoricalNB
from sklearn.model_selection import KFold
from sklearn.metrics import f1_score, accuracy_score
from sklearn.preprocessing import OrdinalEncoder

# 数据加载与类别平衡
df = pd.read_csv('./Datasets/dataset.csv')
grad = df[df['Target']=='Graduate']
drop = df[df['Target']=='Dropout']
grad = grad.sample(n=len(drop), random_state=101)
df = pd.concat([grad, drop], axis=0)

Y = df['Target']
feature_cols = ['Course','Mom_Qualification','Mom_Occupation','Dad_Qualification','Dad_Occupation','Displaced','Age_Enrolled']
k = 10

f1s = []
accuracy = []

for feat in feature_cols:
    # 初始化编码器并基于全数据集拟合,覆盖所有类别
    encoder = OrdinalEncoder()
    singleX = encoder.fit_transform(df[[feat]])
    # 获取该特征的总类别数(用于设置min_categories)
    total_cats = len(encoder.categories_[0])
    
    # 用当前单特征集做KFold拆分
    kf = KFold(n_splits=k, shuffle=True, random_state=101)
    for train_index, test_index in kf.split(singleX):
        trainX = singleX[train_index]
        trainY = Y.iloc[train_index]
        testX = singleX[test_index]
        testY = Y.iloc[test_index]
        
        # 初始化模型,设置min_categories为总类别数避免训练集类别缺失问题
        clf = CategoricalNB(min_categories=total_cats)
        clf.fit(trainX, trainY)
        
        predictions = clf.predict(testX)
        f1s.append(f1_score(testY, predictions, pos_label="Graduate"))
        accuracy.append(accuracy_score(testY, predictions))

补充说明

  • 使用OrdinalEncoder而非LabelEncoder:LabelEncoder仅适用于目标变量,OrdinalEncoder专门处理特征列,且能保持二维数组格式,符合Scikit-learn模型的输入要求。
  • 设置random_state=101:确保KFold拆分结果可复现,便于调试。
  • 全数据集拟合编码器:避免测试集出现训练集未见过的类别,导致模型无法处理。

内容的提问来源于stack exchange,提问作者Nico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 18:00:40