使用HyperOpt优化XGBoost超参数时损失值无变化问题求助
问题原因
- colsample_bytree参数被错误转为整型
你设置的colsample_bytree搜索范围是0.5~1之间的浮点数,但在模型初始化时对该参数用了int()强制转换。Python中0.5到1(不含1)的浮点数转整型结果均为0,会导致每一轮训练时特征采样比例为0,模型无法学习有效特征,所有轮次效果完全一致。 - learning_rate参数未传入模型
搜索空间中定义了learning_rate的搜索范围,但XGBClassifier初始化参数中未加入该字段,模型始终使用默认的0.1学习率,该参数的搜索完全不生效。 - 次要冗余逻辑:
clf.predict()输出的是0/1类别标签,后续pred>0.5的判断没有实际作用,不影响结果但属于多余逻辑。
修复方案
修改objective函数中的模型初始化逻辑,修正上述问题即可,参考代码如下:
def objective(space): clf=xgb.XGBClassifier( n_estimators =space['n_estimators'], max_depth = int(space['max_depth']), gamma = space['gamma'], reg_alpha = int(space['reg_alpha']), reg_lambda = space['reg_lambda'], min_child_weight=int(space['min_child_weight']), colsample_bytree=space['colsample_bytree'], # 移除int()转换 learning_rate=space['learning_rate'], # 新增传入learning_rate参数 seed=space['seed'] ) evaluation = [( X_train, y_train), ( X_test, y_test)] clf.fit(X_train, y_train, eval_set=evaluation, eval_metric="auc", early_stopping_rounds=10,verbose=False) pred = clf.predict(X_test) accuracy = f1_score(y_test, pred) # 移除多余的>0.5判断 print ("SCORE:", accuracy) return {'loss': 1-accuracy, 'status': STATUS_OK }
额外优化建议:你设置的n_estimators为100000数值过大,配合早停机制虽然不会跑完所有轮次,但数据集较大时会拖慢搜索速度,可根据实际情况适当调小该参数。
内容的提问来源于stack exchange,提问作者JHWCHUNG
相关产品推荐
相关产品推荐

