You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

设置enable_categorical=True后仍无法处理分类数据的问题求助

XGBoost启用分类特征报错的解决方法

问题重现

已将数据集中的A、B、C列转为category类型:

cols = ['A','B', 'C']
for i in cols:
        data[i] = data[i].astype('category')

创建XGBoost分类器时已指定enable_categorical=True:

model_xgb = XGBClassifier(**xgb_params, scale_pos_weight=weight, tree_method='hist', enable_categorical=True)

运行时触发如下错误:

ValueError: DataFrame.dtypes for data must be int, float, bool or category. 
When categorical type is supplied, The experimental DMatrix parameter enable_categorical 
must be set to True.  Invalid columns:A: category, B: category, C: category

解决方法

  • 确认XGBoost版本兼容性:enable_categorical是1.3.0版本后引入的实验特性,版本过低会导致参数不生效。执行print(xgboost.__version__)查看版本,低于1.3.0则升级:

    pip install --upgrade xgboost
    
  • 显式构建DMatrix并设置参数:部分场景下XGBClassifier.fit()不会自动传递enable_categorical到底层DMatrix,需手动构建DMatrix并指定参数:

    import xgboost as xgb
    
    # 拆分特征与目标变量
    X = data[['A', 'B', 'C']]
    y = data['y']
    
    # 构建带分类支持的DMatrix
    dtrain = xgb.DMatrix(X, label=y, enable_categorical=True)
    
    # 训练模型
    model_xgb = xgb.train(
        params=xgb_params,
        dtrain=dtrain,
        scale_pos_weight=weight,
        tree_method='hist'
    )
    

    若坚持使用XGBClassifier,可尝试在fit()时为评估集也传入构建好的DMatrix,确保参数全局生效。

  • 校验分类列类型:执行print(data.dtypes)确认A、B、C列确实为category类型,避免后续操作(如缺失值填充、合并数据集)意外将其转为object类型。

  • 禁止重复编码:不要对已转为category的列执行独热编码(如pd.get_dummies),否则会破坏分类特征的原生格式,导致XGBoost无法识别。

内容的提问来源于stack exchange,提问作者Charun Umesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 23:17:25