XGBoost分类:XGBClassifier与xgb.train的结果一致性及用法
使用xgb.train实现分类器的正确方法(二分类/多分类)
二分类场景的处理逻辑
当你用objective="binary:logistic"训练二分类模型时,xgb.train返回的模型调用predict()确实会输出样本属于正类(标签1)的概率值。此时用0.5作为阈值划分类别是完全正确的:概率>0.5归为1类,<0.5归为0类,和XGBClassifier.predict()的内置逻辑完全一致——你的测试代码已经验证了这一点:通过round()处理概率后得到的分类结果,和XGBClassifier.predict()的输出完全匹配,两者的概率值也完全对应。
多分类场景的正确处理(重点:不要用均匀区间划分)
多分类场景下绝对不能用均匀区间(比如3类划分为01/3、1/32/3、2/3~1)来划分,正确的做法是选择对应的多分类目标函数,配合num_class参数(必须指定类别数量):
- 目标函数选
multi:softmax:predict()直接返回样本的类别标签(整数形式),无需额外处理; - 目标函数选
multi:softprob:predict()返回二维概率矩阵(每行对应一个样本,每列对应对应类别的概率),此时要得到类别标签,需要取每行概率最大值对应的索引(比如用np.argmax(preds, axis=1))。
多分类示例代码
import numpy as np import xgboost as xgb # 生成3分类测试数据 data = np.random.rand(50, 10) label = np.random.randint(3, size=50) # 3类标签:0、1、2 dtrain = xgb.DMatrix(data, label=label) # 多分类参数配置:指定num_class,选softmax或softprob param = { 'max_depth':3, 'eta':0.1, 'silent':1, 'tree_method':'hist', 'objective':'multi:softmax', # 换成multi:softprob会返回概率矩阵 'num_class':3, # 必须指定类别数 'seed':42 } num_round = 100 bst = xgb.train(param, dtrain, num_round) # 预测 preds = bst.predict(dtrain) print(preds) # 用softmax时直接输出类别标签:[0 2 1 ...]
你的测试代码验证
你提供的二分类测试代码已经很好地验证了xgb.train和XGBClassifier的一致性:
import numpy as np import xgboost as xgb data = np.random.rand(50,10) # 50 entities, each contains 10 features label = np.random.randint(2, size=50) # binary target dtrain = xgb.DMatrix(data, label=label) param = {'max_depth':3, 'eta':0.1, 'silent':1, 'tree_method':'hist','objective':'binary:logistic', 'seed':42} num_round = 100 # same as number of estimator bst = xgb.train( param, dtrain, num_round) trainres = bst.predict(dtrain) model = xgb.XGBClassifier(n_estimators=100, objective='binary:logistic', tree_method='hist', eta=0.1, max_depth=3, enable_categorical=True, seed=42) model = model.fit(data,label) fitres = model.predict(data) # Compare classification print(all([round(x) for x in trainres] == fitres)) # Compare probabilities. Predict proba gives prob for class 0 and 1, so take x[1] print(all([x[1] for x in model.predict_proba(data)] == trainres))
输出结果:
True True
总结使用xgb.train做分类的步骤
- 将训练/测试数据转换为
DMatrix格式; - 根据分类类型配置参数:
- 二分类:设置
objective="binary:logistic",可选添加eval_metric(如error); - 多分类:设置
objective为multi:softmax或multi:softprob,必须指定num_class;
- 二分类:设置
- 调用
xgb.train()训练模型,传入参数、DMatrix和迭代次数; - 预测与分类:
- 二分类:用0.5阈值对概率做划分(或
np.round())得到类别; - 多分类(softmax):直接使用
predict()结果作为类别; - 多分类(softprob):对概率矩阵取每行最大值的索引得到类别。
- 二分类:用0.5阈值对概率做划分(或
内容的提问来源于stack exchange,提问作者Lim Kaizhuo
相关产品推荐
相关产品推荐

