XGBoost多分类遇ValueError:类别数与target_names不匹配及变量类型疑问
问题
编写XGBoost多分类代码时触发错误:
ValueError: Number of classes does not match size of target_names for a confusion matrix
代码详情:
y = df['weather type'] # y是包含11个唯一值0,1,2,3,4,5,6,7,8,9,10的object类型数组: # array([10, 10, 10, ..., 10, 10, 9]) # 编码目标变量(天气类型)为数值型——不确定这一步是否必要,好像搞乱了标签? y_le = LabelEncoder() y = y_le.fit_transform(y) # y_le.classes_的唯一值是'0', '1', '10', '11', '12', '2', '3', '5', '6', '7', '8' # y_val的唯一值是0, 1, 3, 4, 5, 6, 7, 8, 9, 10 # 初始化XGBoost分类器 xgb_model = xgb.XGBClassifier(objective='multi:softmax', num_class=len(le.classes_)) # 训练模型 xgb_model.fit(X_train, y_train) # 在验证集上预测 y_pred_val = grid_search.predict(X_val) # 评估模型:打印分类报告和混淆矩阵 print("\nClassification Report:\n", classification_report(y_val, y_pred_val, target_names=y_le.classes_))
完整错误栈:
-------------------------------------------------------------------------- ValueError Traceback (most recent call last) Cell In[292], line 6 2 y_pred_val = xgb_model.predict(X_val) 4 # Evaluate the model 5 # Print classification report and confusion matrix ----> 6 print("\nClassification Report:\n", classification_report(y_val, y_pred_val, target_names=y_le.classes_)) 7 #print("\nClassification Report:\n", classification_report(y_val, y_pred_val, labels=range(len(y_le.classes_)), target_names=y_le.classes_)) File ~\anaconda3\lib\site-packages\sklearn\metrics\_classification.py:2332, in classification_report(y_true, y_pred, labels, target_names, sample_weight, digits, output_dict, zero_division) 2326 warnings.warn( 2327 "labels size, {0}, does not match size of target_names, {1}".format( 2328 len(labels), len(target_names) 2330 ) 2331 else: -> 2332 raise ValueError( 2333 "Number of classes, {0}, does not match size of " 2334 "target_names, {1}. Try specifying the labels " 2335 "parameter".format(len(labels), len(target_names)) 2336 ) 2337 if target_names is None: 2338 target_names = ["%s" % l for l in labels] ValueError: Number of classes, 10, does not match size of target_names, 11. Try specifying the labels parameter
疑问:
- 已设置
target_names=y_le.classes_仍报错,如何修复? - 目标变量
weather_type是object类型,是否需要转为数值型用于XGBoost多分类?
解决方案
1. 修复分类报告的类数量不匹配错误
错误核心是:y_val仅包含10个类别,但y_le.classes_有11个类别(存在'11'、'12'这类验证集里没有的类别),导致真实类别数和目标名称数不一致。
两种修复方式:
- 指定
labels参数:明确告诉函数所有需要统计的类别编码,确保和target_names长度一致:
# 生成所有类别的编码标签 all_labels = y_le.transform(y_le.classes_) print("\nClassification Report:\n", classification_report( y_val, y_pred_val, labels=all_labels, target_names=y_le.classes_ ))
这样即使验证集没有某些类别,函数也会保留对应列,避免数量不匹配。
- 清理数据一致性:看你原始注释,y的原始值是0-10,说明
'11'、'12'是异常值,应该先过滤掉这些不存在的类别,再做编码,从根源避免类别数量不一致的问题。
2. 关于目标变量是否需要转数值型
XGBoost的XGBClassifier不能直接处理object类型的字符串标签,必须转成数值型,但要选对编码方式:
- 你的
weather type是object类型的数字字符串,直接用astype(int)转成整数更合适,避免LabelEncoder按字典序排序导致的混乱(比如你现在的y_le.classes_里'10'排在'2'前面,编码后'10'对应1,'2'对应5,会干扰模型对类别的认知)。 - 代码示例:
y = df['weather type'].astype(int)
如果是非数字的类别名称,再用LabelEncoder或OrdinalEncoder编码。
内容的提问来源于stack exchange,提问作者Bluetail
相关产品推荐
相关产品推荐

