You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于LGBM(Sklearn API)与Optuna:早停剪枝及自定义评估指标问题

LGBM(Sklearn API)结合Optuna调优的核心问题解答

1. Pruning(剪枝)与Early Stopping(早停)是同一机制吗?

不是,二者是独立控制不同流程的机制:

  • Early Stopping:属于LightGBM的训练控制逻辑,作用是在单个模型训练过程中,当验证集指标不再提升(或下降)时提前停止迭代,避免过拟合、节省训练时间。代码中early_stopping(100)表示连续100轮迭代指标无优化就停止。
  • Pruning(Optuna剪枝):属于Optuna的超参数调优逻辑,作用是提前终止表现极差的超参数trial(即某一组超参数对应的模型),避免在无效参数上浪费资源。比如如果某组参数在训练早期就比当前最优trial差很多,Optuna会直接砍掉这个trial,不用等到模型训练完成。

代码里的两个回调独立生效:LightGBMPruningCallback负责Optuna的剪枝逻辑,early_stopping负责单个模型的早停逻辑。

2. 不平衡目标下如何替换默认指标为自定义评估指标(如平衡准确率)?

LightGBM支持自定义评估函数,需按以下步骤修改:

步骤1:定义符合要求的自定义评估函数

平衡准确率的评估函数需要返回指标名称、指标值、是否需要最大化(True表示指标越大越好):

from sklearn.metrics import balanced_accuracy_score
import numpy as np

def balanced_acc_eval(y_true, y_pred):
    # LightGBM自定义评估函数格式:(名称, 指标值, 是否更大更好)
    y_pred = np.round(y_pred)  # 将LGBM输出的概率转为类别标签
    score = balanced_accuracy_score(y_true, y_pred)
    return "balanced_acc", score, True

步骤2:修改调参逻辑中的相关参数

替换eval_metric为自定义函数,同时更新LightGBMPruningCallback的监控指标,确保Optuna的优化方向与指标匹配(用1 - 平衡准确率转为最小化问题):

def objective(trial, X, y):
    param_grid = {
        "n_estimators": trial.suggest_categorical("n_estimators", [999999]),
        "learning_rate": trial.suggest_float("learning_rate", 0.01, 0.3),
        "num_leaves": trial.suggest_int("num_leaves", 20, 3000, step=20),
        "max_depth": trial.suggest_int("max_depth", 3, 12),
        "min_data_in_leaf": trial.suggest_int("min_data_in_leaf", 200, 10000, step=100),
        "lambda_l1": trial.suggest_int("lambda_l1", 0, 100, step=5),
        "lambda_l2": trial.suggest_int("lambda_l2", 0, 100, step=5),
        "min_gain_to_split": trial.suggest_float("min_gain_to_split", 0, 15),
        "bagging_fraction": trial.suggest_float("bagging_fraction", 0.2, 0.95, step=0.1),
        "bagging_freq": trial.suggest_categorical("bagging_freq", [1]),
        "feature_fraction": trial.suggest_float("feature_fraction", 0.2, 0.95, step=0.1),
    }

    cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=1121218)
    cv_scores = np.empty(5)
    best_iterations = []  # 记录每个fold的最优迭代次数,供后续问题3使用

    for idx, (train_idx, test_idx) in enumerate(cv.split(X, y)):
        X_train, X_test = X.iloc[train_idx], X.iloc[test_idx]
        y_train, y_test = y.iloc[train_idx], y.iloc[test_idx]

        model = LGBMClassifier(
            objective="binary",
            **param_grid,
            n_jobs=-1,
            scale_pos_weight=len(y_train) / y_train.sum()
        )
        
        model.fit( 
            X_train,
            y_train,
            eval_set=[(X_test, y_test)],
            eval_metric=balanced_acc_eval,  # 替换为自定义评估函数
            callbacks=[
                LightGBMPruningCallback(trial, "balanced_acc"),  # 监控自定义指标
                early_stopping(100, verbose=False)
            ], 
        )
        preds = model.predict(X_test)
        cv_scores[idx] = balanced_accuracy_score(y_test, preds)
        best_iterations.append(model.best_iteration_)  # 记录当前fold的最优迭代次数
    
    loss = 1 - np.nanmedian(cv_scores)
    # 将最优迭代次数的中位数存入trial属性,方便后续调用
    trial.set_user_attr("median_best_iter", int(np.median(best_iterations)))
    return loss

3. 如何利用剪枝后的最优n_estimators训练最终模型?

调参时n_estimators设为极大值,每个fold训练后的模型会通过早停得到best_iteration_属性(该组参数下的最优迭代次数),按以下步骤复用该值:

步骤1:运行调优并提取最优参数与迭代次数

study = optuna.create_study(direction="minimize", study_name="LGBM Classifier")
func = lambda trial: objective(trial, X_train, y_train)
study.optimize(func, n_trials=50)  # 根据需求调整trial数量

# 复制最优参数并替换n_estimators
best_params = study.best_params.copy()
best_n_estimators = study.best_trial.user_attrs["median_best_iter"]
best_params["n_estimators"] = best_n_estimators

步骤2:训练最终模型

final_model = LGBMClassifier(
    objective="binary",
    **best_params,
    n_jobs=-1,
    scale_pos_weight=len(y_train) / y_train.sum()
)

# 可选:最终训练时拆分验证集再次用早停确认最优迭代
from sklearn.model_selection import train_test_split
X_train_final, X_val, y_train_final, y_val = train_test_split(
    X_train, y_train, test_size=0.1, stratify=y_train, random_state=42
)
final_model.fit(
    X_train_final,
    y_train_final,
    eval_set=[(X_val, y_val)],
    eval_metric=balanced_acc_eval,
    callbacks=[early_stopping(100, verbose=False)]
)
# 若使用早停,可直接用final_model.best_iteration_作为最终迭代次数

内容的提问来源于stack exchange,提问作者Kjetil Haukås

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 10:50:41