You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Hyperopt优化LightGBM时MLflow log_params参数重复报错问题

问题解决:Hyperopt优化LightGBM时MLflow参数记录冲突

错误原因

你把Hyperopt的所有超参数迭代都放在同一个MLflow主Run里,而MLflow的规则是:同一个Run中,同一个参数键只能记录一次值。当第二次迭代的参数(比如colsample_bytree)和第一次值不同时,就触发了INVALID_PARAMETER_VALUE错误。

推荐解决方案:给每个Hyperopt迭代创建独立嵌套子Run

每个超参数组合对应一个独立的MLflow子Run,参数和指标可以独立记录,还能在MLflow UI中清晰对比所有迭代结果。

修改后的完整代码(补全了缺失的numpy导入):

from sklearn.metrics import f1_score
import lightgbm as lgbm
import hyperopt
import numpy as np  # 补全缺失的numpy导入
from hyperopt import fmin, tpe, hp, STATUS_OK, space_eval, Trials, SparkTrials
from hyperopt.pyll.base import scope 
import mlflow


lgbm_space = {
        'boosting_type': hp.choice('boosting_type', ['gbdt', 'dart', 'goss']),
        'n_estimators': hp.choice('n_estimators', np.arange(400, 1000, 50, dtype=int)), 
        'learning_rate' : hp.quniform('learning_rate', 0.02, 0.5, 0.02), 
        'max_depth': scope.int(hp.quniform('max_depth', 2, 16, 1)),
        'num_leaves': hp.choice("num_leaves", np.arange(10, 80, 5, dtype=int)),
        'colsample_bytree': hp.uniform('colsample_bytree', 0.7, 1.0),
        'subsample': hp.uniform('subsample', 0.7, 1.0), 
        'min_child_samples': hp.choice('min_child_samples', np.arange(10, 50, 5, dtype=int))
}

search_space = lgbm_space
run_name = "run_optimization" 
max_eval = 100

#define objective function
def objective(search_space):
    # 为每个迭代创建嵌套子Run
    with mlflow.start_run(nested=True):
        model = lgbm.LGBMClassifier( **search_space, class_weight='balanced', n_jobs=-1, random_state=123 )      
        model.fit(X_train, y_train,            
               eval_set= [ ( X_val, y_val) ], 
               early_stopping_rounds= 10, 
               verbose=False)    
        y_pred = model.predict_proba(X_val)[:,1]   
        f1 = f1_score(y_val, (y_pred>0.5).astype(int) )
        mlflow.log_metric('f1 score', f1)
        mlflow.log_params(search_space)
        # 可选:记录训练好的模型
        mlflow.lightgbm.log_model(model, "model")
        score = 1 - f1
    
        return {'loss': score, 'status': STATUS_OK, 'model': model, 'params': search_space}

spark_trials = Trials()
with mlflow.start_run(run_name = run_name):
    best_params = hyperopt.fmin(
                    fn = objective,
                    space = search_space,
                    algo = tpe.suggest,
                    max_evals = max_eval, 
                    trials = spark_trials )

关键修改点

  • 在objective函数内部添加mlflow.start_run(nested=True):启动嵌套子Run,每个迭代的参数和指标都会被记录到独立的子Run中,不会和其他迭代冲突。
  • 子Run会自动关联到外层的主Run,在MLflow UI中可以展开主Run查看所有迭代的细节。

备选方案(不推荐):同一Run中批量记录所有迭代结果

如果非要在同一个主Run中记录所有结果,可以把每次的参数和指标存储到Trials对象里,最后统一用mlflow.log_dict记录,但这种方式在UI中无法直观查看单个迭代的结果:

# 修改objective函数,不直接记录MLflow,而是把结果存在trials里
def objective(search_space):
    model = lgbm.LGBMClassifier( **search_space, class_weight='balanced', n_jobs=-1, random_state=123 )      
    model.fit(X_train, y_train,            
           eval_set= [ ( X_val, y_val) ], 
           early_stopping_rounds= 10, 
           verbose=False)    
    y_pred = model.predict_proba(X_val)[:,1]   
    f1 = f1_score(y_val, (y_pred>0.5).astype(int) )
    score = 1 - f1
    
    # 把参数和指标返回,存在trials中
    return {'loss': score, 'status': STATUS_OK, 'model': model, 'params': search_space, 'f1': f1}

# 主Run中统一记录所有迭代结果
spark_trials = Trials()
with mlflow.start_run(run_name = run_name):
    best_params = hyperopt.fmin(
                    fn = objective,
                    space = search_space,
                    algo = tpe.suggest,
                    max_evals = max_eval, 
                    trials = spark_trials )
    # 遍历trials,把所有迭代的参数和指标记录为字典
    all_iterations = []
    for trial in spark_trials.trials:
        all_iterations.append({
            'params': trial['result']['params'],
            'f1_score': trial['result']['f1'],
            'loss': trial['result']['loss']
        })
    mlflow.log_dict(all_iterations, "all_hyperopt_iterations.json")

内容的提问来源于stack exchange,提问作者zesla

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:34:52