You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

含分类变量的XGBoost模型导出失败求助:特征类型不匹配报错

问题分析与解决建议

问题根源

报错Check failed: is_categorical: A in feature map is numerical but tree node is categorical的核心原因:

  • 代码存在数据赋值错误,导致特征列数据被意外覆盖;
  • XGBoost旧版本中,启用enable_categorical=True训练后,trees_to_dataframe()方法存在类型识别bug——模型内部将A识别为分类特征,但导出时特征映射误标记为数值型,引发类型不匹配。

代码中的明显错误

你代码里这一行逻辑错误:

X[['B', 'C']] = X[['A', 'B']].astype('float')

这行把原A列的分类数据赋值给B列,原B列数据赋值给C列,完全覆盖了B、C列的原始数值数据,必须修正为:

X[['B', 'C']] = X[['B', 'C']].astype('float')

解决步骤

1. 修正数据赋值错误

先修复上述列赋值错误,确保训练数据的特征列数据正确。

2. 显式指定特征类型

创建DMatrix时,通过feature_types参数明确声明每个特征的类型,强制模型识别分类特征:

import numpy as np
import pandas as pd
import xgboost as xgb

X1 = [0, 2, 3, 1, 4, 5]
X2 = np.random.rand(6)
X3 = np.random.rand(6)*2
X = pd.DataFrame(data=np.column_stack((X1, X2, X3)), columns=['A', 'B', 'C'])
X['A'] = X['A'].astype('category')
X[['B', 'C']] = X[['B', 'C']].astype('float')  # 修正后的赋值
y = pd.DataFrame(data=np.random.rand(6)).astype('float')

XGBParams = {'booster': 'gbtree'}
# 显式指定特征类型:'c'代表分类,'q'代表数值
feature_types = ['c', 'q', 'q']
d = xgb.DMatrix(X, label=y, missing=np.NaN, enable_categorical=True, feature_types=feature_types)
model = xgb.train(XGBParams, d, num_boost_round=20, verbose_eval=True)
print(model.trees_to_dataframe())

3. 升级XGBoost版本

该类型不匹配报错是XGBoost 1.7.0版本之前的已知bug,升级到最新稳定版可彻底解决:

pip install --upgrade xgboost

4. 替代导出方案(若上述方法无效)

如果升级后仍无法使用trees_to_dataframe(),可以先将模型导出为JSON,再手动转换为DataFrame:

# 导出模型为JSON字符串
model_json_str = model.get_booster().dump_model(dump_format='json')
# 解析JSON并转为DataFrame
import json
tree_data = json.loads(model_json_str)
df = pd.json_normalize(tree_data)
print(df)

内容的提问来源于stack exchange,提问作者CaptBarnacles

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 14:53:10