You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于MAE和RMSE用Optuna调优LGBM回归器时遇ValueError

问题排查与解决方案

错误根源

你遇到的ValueError源于两个核心问题:

  1. LightGBMPruningCallback不支持多指标列表:该回调的metric参数仅接受单个字符串(如'l1'或'rmse'),但你传入了列表['l1', 'rmse']。回调会尝试查找名为['l1', 'rmse']的指标,而实际LightGBM返回的是分开的'l1'和'rmse'指标条目,因此匹配失败。
  2. 多指标场景下early_stopping未指定监控目标:当eval_metric传入多个指标时,early_stopping回调需要通过monitor参数明确指定基于哪个指标触发早停,否则会出现行为异常。

修正方案

针对多目标优化场景,我们可以通过手动报告指标+触发剪枝的方式替代原有的LightGBMPruningCallback,同时为early_stopping指定明确的监控指标。以下是修正后的完整代码:

# Get categorical features
cat_features = df.select_dtypes(include='category').columns.to_list()

def objective(trial, X, y):
    param_grid = {
        'objective': 'regression',
        'n_estimators': trial.suggest_int('n_estimators', 100, 10000, step=100),
        'learning_rate': trial.suggest_float('learning_rate', 0.01, 0.3),
        'num_leaves': trial.suggest_int('num_leaves', 20, 3000, step=20),
        'max_depth': trial.suggest_int('max_depth', 3, 12),
        'min_data_in_leaf': trial.suggest_int('min_data_in_leaf', 200, 10000, step=100),
        'max_bin': trial.suggest_int('max_bin', 200, 300),
        'lambda_l1': trial.suggest_int('lambda_l1', 0, 100, step=5),
        'lambda_l2': trial.suggest_int('lambda_l2', 0, 100, step=5),
        'min_gain_to_split': trial.suggest_float('min_gain_to_split', 0, 15),
        'bagging_fraction': trial.suggest_float('bagging_fraction', 0.2, 0.95, step=0.1),
        'feature_fraction': trial.suggest_float('feature_fraction', 0.2, 0.95, step=0.1)
    }

    cv = KFold(n_splits=5, shuffle=True, random_state=42)

    cv_scores_mae = []
    cv_scores_rmse = []

    for idx, (train_idx, test_idx) in enumerate(cv.split(X, y)):
        X_train, X_test = X.iloc[train_idx], X.iloc[test_idx]
        y_train, y_test = y[train_idx], y[test_idx]

        model = lgbm.LGBMRegressor(**param_grid)

        model.fit(
            X_train, y_train,
            eval_set=[(X_test, y_test)],
            eval_metric=['l1', 'rmse'],
            categorical_feature=cat_features,
            # 指定早停监控的指标(这里选择rmse,可根据需求换成l1)
            callbacks=[lgbm.early_stopping(50, monitor='valid_0_rmse')],
            verbose=False
        )

        y_pred = model.predict(X_test)

        # Calculate the evaluation metrics
        mae = mean_absolute_error(y_test, y_pred)
        rmse = np.sqrt(mean_squared_error(y_test, y_pred))

        cv_scores_mae.append(mae)
        cv_scores_rmse.append(rmse)

        # 手动向Optuna报告当前fold的两个指标
        trial.report(mae, idx)
        trial.report(rmse, idx)
        # 检查是否需要剪枝
        if trial.should_prune():
            raise optuna.TrialPruned()

    return np.mean(cv_scores_mae), np.mean(cv_scores_rmse)

study = optuna.create_study(directions=['minimize', 'minimize'], study_name="LGBM Regressor")
func = lambda trial: objective(trial, df.drop(columns='price'), df['price'])
study.optimize(func, n_trials=20, show_progress_bar=True)

关键修改说明

  • 移除了LightGBMPruningCallback,改用trial.report()分别报告MAE和RMSE指标,再通过trial.should_prune()触发剪枝逻辑,适配多目标优化场景。
  • 为early_stopping回调添加monitor='valid_0_rmse'参数,明确基于RMSE指标触发早停(若需基于MAE,可改为'valid_0_l1')。
  • 保留了原有的交叉验证和多指标返回逻辑,确保Optuna能同时优化MAE和RMSE两个目标。

内容的提问来源于stack exchange,提问作者Roman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 00:14:51