基于AUC早停的LightGBM二分类器:目标函数是Log Loss还是AUC?
问题
我用LightGBM梯度提升模型做二分类任务,按照LightGBM参数规范,二分类的目标函数是log loss(交叉熵),这是模型优化时采用的目标函数。但我设置了若验证集AUC在10轮内无提升则停止训练,想问算法的目标函数会不会变成基于AUC的函数,还是仍保持log loss,仅在验证集上额外计算AUC来判断是否早停?
参考代码如下:
model = lightgbm.LGBMClassifier(n_estimators=500, learning_rate=0.01) fit_params={"early_stopping_rounds":30, "eval_metric" : 'auc', "eval_set" : [(X_test,y_test)], 'eval_names': ['valid'], #'callbacks': [lgb.reset_parameter(learning_rate=learning_rate_010_decay_power_099)], 'verbose': 100, 'categorical_feature': 'auto'} params = { 'boosting': ['gbdt'], 'objective': ['binary'], 'num_leaves': sp_randint(20, 63), # adjust to inc AUC 'max_depth': sp_randint(3, 8), 'min_child_samples': sp_randint(100, 500), 'min_child_weight': [1e-5, 1e-3, 1e-2, 1e-1, 1, 1e1, 1e2, 1e3, 1e4], 'subsample': sp_uniform(loc=0.2, scale=0.8), 'colsample_bytree': sp_uniform(loc=0.4, scale=0.6), 'reg_alpha': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], 'reg_lambda': [0, 1e-1, 1, 5, 10, 20, 50, 100], 'is_unbalance': ['true'], 'feature_fraction': np.arange(0.3, 0.6, 0.05), # model trained faster AND prevents overfitting 'bagging_fraction': np.arange(0.3, 0.6, 0.05), # similar benefits as above 'bagging_freq': np.arange(10, 30, 5), # after every 20 iterations, lgb will randomly select 50% of observations and use it for next 20 iterations } RSCV = RandomizedSearchCV(model, params, scoring= 'roc_auc', cv=3, n_iter=10, refit=True, random_state=42, verbose=True) RSCV.fit(X_train, y_train, **fit_params)
回答
模型的目标函数仍然是log loss,不会因为设置了基于AUC的早停逻辑而改变。
具体细节:
- 训练全程,模型始终以最小化log loss为优化核心,每一轮迭代的参数更新都是围绕这个目标展开的。
eval_metric='auc'只是指定在验证集上额外计算AUC作为监控指标,用来观察模型的泛化表现;early_stopping_rounds则是依据这个AUC的变化来触发早停——当验证集AUC连续指定轮数没有提升时,终止训练以避免过拟合。- 简言之,目标函数是模型优化的“方向”,早停用的AUC是判断优化何时该停止的“监控器”,二者互不干扰。
内容的提问来源于stack exchange,提问作者Jim R.
相关产品推荐
相关产品推荐

