You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XGBoost分类任务中目标为字符串类别时的报错问题咨询

问题原因及解决方法

你遇到的报错核心是标签编码逻辑和模型输入输出不匹配,具体问题和修正方案如下:

核心报错原因

你虽然初始化了LabelEncoder并对原始标签做了拟合,但没有将编码后的数值标签替换掉原始字符串标签用于模型训练:

  • 拆分训练测试集、调用model.fit()时用的还是原始的y_train(值为Good/Bad字符串)
  • 模型训练时接收的是字符串标签,预测输出的自然也是字符串,后续强行对字符串做int()转换就会触发报错

完整修正代码

import xgboost as xgb
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder 
from xgboost import XGBClassifier
from sklearn.metrics import accuracy_score
import pandas as pd

# 路径注意转义,用r开头的原始字符串避免转义字符报错
data = r'C:\me\my_table.csv'
df = pd.read_csv(data)

cols_to_drop = ['_id']
df.drop(cols_to_drop, axis=1, inplace=True)

X = df.drop('Rank', axis=1)
y = df['Rank']

# 标签编码:将字符串标签转为数值,且直接替换原始y用于后续流程
lc = LabelEncoder() 
y = lc.fit_transform(y)

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state=5)

model = XGBClassifier() 
model.fit(X_train, y_train)

y_pred = model.predict(X_test) 
# 此时y_pred已经是0/1的数值类型,无需额外转int/round
accuracy = accuracy_score(y_test, y_pred)
print(f"模型准确率:{accuracy:.2f}")

# 如果需要把预测结果转回原始字符串标签,可以用 inverse_transform
pred_labels = lc.inverse_transform(y_pred)

额外注意事项

  • 你数据集中存在大量NaN的字段,XGBoost原生支持缺失值处理,无需额外填充;但如果某字段缺失比例超过80%,建议直接删除该字段避免引入噪声
  • 如果后续类别数量超过2类,上述标签编码逻辑依然适用,无需额外修改

内容的提问来源于stack exchange,提问作者futuredataengineer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 04:06:06