You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

XGBoost的base_score如何计算?为何与目标均值存在差异?

关于XGBoost 2.0+自动计算base_score的逻辑疑问

从XGBoost 2.0版本起,初始化estimator时若未指定base_score参数,该参数会自动计算。我原本以为它直接使用目标变量的均值,但实际情况并非如此:

import json
import shap # 仅用于获取数据集
import xgboost as xgb

print('shap.__version__:',shap.__version__)
print('xgb.__version__:',xgb.__version__)
print()

X, y = shap.datasets.adult()

estimator = xgb.XGBClassifier(
    objective='binary:logistic',
    n_estimators=200)

estimator.fit(X,y)

print('y.mean():',y.mean())
print("base_score值:", float(json.loads(estimator.get_booster().save_config())['learner']['learner_model_param']['base_score']))

运行输出:

shap.__version__: 0.46.0
xgb.__version__: 2.1.0

y.mean(): 0.2408095574460244
base_score值: 0.26177529

二者差异明显,绝非舍入误差导致。那么base_score到底是怎么计算出来的?我查到有相关的代码提交,但没法明确具体的计算逻辑。

内容的提问来源于stack exchange,提问作者dwolfeu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 09:35:55