XGBoost回归贝叶斯超参数优化:无法调优n_estimators?
问题原因及解决方法
核心原因
XGBoost的原生低级API(比如xgb.cv)和scikit-learn封装的API(比如XGBRegressor)对迭代次数的参数定义存在差异:
n_estimators是scikit-learn接口特有的参数,用于指定模型树的数量;- 而
xgb.cv这类原生API,是通过num_boost_round参数控制迭代次数,它完全不识别n_estimators,所以将n_estimators放入params字典传给xgb.cv时,会被提示“未使用”。
你的带交叉验证代码中,既手动设置了num_boost_round=100,又将n_estimators塞进params,属于无效操作——原生API只会读取num_boost_round的值,直接忽略n_estimators。
修复代码
将n_estimators从params字典中移除,转而作为num_boost_round的参数传入xgb.cv,修改后的xgb_cv函数如下:
from bayes_opt import BayesianOptimization def xgb_cv(max_depth, learning_rate, subsample, colsample_bytree, lambd, alpha, min_child_weight, gamma, scale_pos_weight, n_estimators): params = { 'objective': 'reg:squarederror', 'max_depth': int(max_depth), 'learning_rate': learning_rate, 'subsample': subsample, 'colsample_bytree': colsample_bytree, 'lambda': lambd, 'alpha': alpha, 'min_child_weight': min_child_weight, 'gamma': gamma, 'scale_pos_weight': scale_pos_weight # 移除n_estimators参数 } dtrain = xgb.DMatrix(X_train, label = y_train) # 将n_estimators转为int后传给num_boost_round cv_result = xgb.cv(params, dtrain, num_boost_round=int(n_estimators), early_stopping_rounds=10, nfold=10, metrics='error') return -cv_result['test-error-mean'].iloc[-1]
修改后,n_estimators会被正确用于控制交叉验证时的树数量,警告信息也会消失。
内容的提问来源于stack exchange,提问作者Yujie Liu
相关产品推荐
相关产品推荐

