You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XGBoost多分类遇ValueError:类别数与target_names不匹配及变量类型疑问

问题

编写XGBoost多分类代码时触发错误:

ValueError: Number of classes does not match size of target_names for a confusion matrix

代码详情:

y = df['weather type'] 
# y是包含11个唯一值0,1,2,3,4,5,6,7,8,9,10的object类型数组:
# array([10, 10, 10, ..., 10, 10,  9])

# 编码目标变量(天气类型)为数值型——不确定这一步是否必要,好像搞乱了标签?
y_le = LabelEncoder()
y = y_le.fit_transform(y)

# y_le.classes_的唯一值是'0', '1', '10', '11', '12', '2', '3', '5', '6', '7', '8'
# y_val的唯一值是0, 1, 3, 4, 5, 6, 7, 8, 9, 10

# 初始化XGBoost分类器
xgb_model = xgb.XGBClassifier(objective='multi:softmax', num_class=len(le.classes_))

# 训练模型
xgb_model.fit(X_train, y_train)

# 在验证集上预测
y_pred_val = grid_search.predict(X_val)

# 评估模型:打印分类报告和混淆矩阵
print("\nClassification Report:\n", classification_report(y_val, y_pred_val, target_names=y_le.classes_))

完整错误栈:

--------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Cell In[292], line 6
      2 y_pred_val = xgb_model.predict(X_val)
      4 # Evaluate the model
      5 # Print classification report and confusion matrix
----> 6 print("\nClassification Report:\n", classification_report(y_val, y_pred_val, target_names=y_le.classes_))
      7 #print("\nClassification Report:\n", classification_report(y_val, y_pred_val, labels=range(len(y_le.classes_)), target_names=y_le.classes_))

File ~\anaconda3\lib\site-packages\sklearn\metrics\_classification.py:2332, in classification_report(y_true, y_pred, labels, target_names, sample_weight, digits, output_dict, zero_division)
   2326         warnings.warn(
   2327             "labels size, {0}, does not match size of target_names, {1}".format(
   2328                 len(labels), len(target_names)
   2330         )
   2331     else:
-> 2332         raise ValueError(
   2333             "Number of classes, {0}, does not match size of "
   2334             "target_names, {1}. Try specifying the labels "
   2335             "parameter".format(len(labels), len(target_names))
   2336         )
   2337 if target_names is None:
   2338     target_names = ["%s" % l for l in labels]

ValueError: Number of classes, 10, does not match size of target_names, 11. Try specifying the labels parameter

疑问:

  1. 已设置target_names=y_le.classes_仍报错,如何修复?
  2. 目标变量weather_type是object类型,是否需要转为数值型用于XGBoost多分类?

解决方案

1. 修复分类报告的类数量不匹配错误

错误核心是:y_val仅包含10个类别,但y_le.classes_有11个类别(存在'11'、'12'这类验证集里没有的类别),导致真实类别数和目标名称数不一致。

两种修复方式:

  • 指定labels参数:明确告诉函数所有需要统计的类别编码,确保和target_names长度一致:
# 生成所有类别的编码标签
all_labels = y_le.transform(y_le.classes_)
print("\nClassification Report:\n", classification_report(
    y_val, 
    y_pred_val, 
    labels=all_labels, 
    target_names=y_le.classes_
))

这样即使验证集没有某些类别,函数也会保留对应列,避免数量不匹配。

  • 清理数据一致性:看你原始注释,y的原始值是0-10,说明'11'、'12'是异常值,应该先过滤掉这些不存在的类别,再做编码,从根源避免类别数量不一致的问题。

2. 关于目标变量是否需要转数值型

XGBoost的XGBClassifier不能直接处理object类型的字符串标签,必须转成数值型,但要选对编码方式:

  • 你的weather type是object类型的数字字符串,直接用astype(int)转成整数更合适,避免LabelEncoder按字典序排序导致的混乱(比如你现在的y_le.classes_里'10'排在'2'前面,编码后'10'对应1,'2'对应5,会干扰模型对类别的认知)。
  • 代码示例:
y = df['weather type'].astype(int)

如果是非数字的类别名称,再用LabelEncoder或OrdinalEncoder编码。


内容的提问来源于stack exchange,提问作者Bluetail

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 11:13:13