You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Optuna调优GPU版CatBoost回归器时进程崩溃问题

问题:Optuna结合GPU调优CatBoost时出现TerminatedWorkerError

背景信息

使用Optuna结合本地NVIDIA GeForce RTX 3060 GPU对CatBoost回归器做超参数调优,数据集规模为120k条训练样本、16个特征(含分类特征)。单独运行CatBoost GPU版本一切正常,但结合Optuna运行时出现进程终止错误;尝试CPU运行速度过慢,添加os.environ['CUDA_VISIBLE_DEVICES'] = '0'并移除n_jobs=-1后问题仍存在。

使用版本:CatBoost 1.2.5,Optuna 4.0.0

运行代码

RANDOM_SEED = 2906
n_trials = 100

def objective_cat(trial):
  params = {'iterations': trial.suggest_int('iterations', 100, 2000),
            'learning_rate': trial.suggest_float('learning_rate', 1e-5, 1e-1, log=True),
            'depth': trial.suggest_int('max_depth', 2, 16),
            'bagging_temperature': trial.suggest_uniform('bagging_temperature', 0, 1),
            'l2_leaf_reg': trial.suggest_float('l2_leaf_reg', 1e-5, 100, log=True),
            'random_seed': RANDOM_SEED,
            'loss_function': 'RMSE',
            'task_type': 'GPU'
            }

  cat_model = CatBoostRegressor(**params, cat_features=cat_features)
  cv = KFold(
      n_splits=10,
      shuffle=True,
      random_state=RANDOM_SEED
      )
  score = cross_val_score(
      cat_model, X_train, y_train, cv=cv, n_jobs=-1,
      scoring='neg_root_mean_squared_error'
      )
  return -np.mean(score)

study_cat = optuna.create_study(direction='minimize',
                                 sampler=optuna.samplers.TPESampler(
                                     seed=RANDOM_SEED,
                                     n_startup_trials=20,
                                     multivariate=True
                                     ),
                                 pruner=optuna.pruners.HyperbandPruner(
                                     min_resource=1,
                                     max_resource=100,
                                     reduction_factor=3
                                     ),
                                 study_name=f'optuna_cat')
optuna.logging.set_verbosity(optuna.logging.CRITICAL)
study_cat.optimize(
    objective_cat,
    n_trials=n_trials,
    n_jobs=-1,
    show_progress_bar=True
    )
best_params_cat = study_cat.best_params

错误信息

TerminatedWorkerError                     Traceback (most recent call last)
Cell In[70], line 39
     26 study_cat = optuna.create_study(direction='minimize',
     27                                  sampler=optuna.samplers.TPESampler(
     28                                      seed=RANDOM_SEED,
   (...)
     36                                      ),
     37                                  study_name=f'optuna_cat')
     38 optuna.logging.set_verbosity(optuna.logging.CRITICAL)
---> 39 study_cat.optimize(
     40     objective_cat,
     41     n_trials=n_trials,
     42     n_jobs=-1,
     43     show_progress_bar=True
...
    764     return self._result
    765 finally:

TerminatedWorkerError: A worker process managed by the executor was unexpectedly terminated. This could be caused by a segmentation fault while calling the function or by an excessive memory usage causing the Operating System to kill the worker.

解决方案

1. 限制并行数量+关闭交叉验证并行

RTX3060显存有限(通常12G),n_jobs=-1会启动与CPU核心数一致的trial进程,每个进程都占用GPU显存,直接导致显存耗尽被系统终止;同时交叉验证的n_jobs=-1会进一步加剧资源冲突。

修改代码:

# 交叉验证部分:将n_jobs改为1
score = cross_val_score(
    cat_model, X_train, y_train, cv=cv, n_jobs=1,
    scoring='neg_root_mean_squared_error'
)

# Optuna优化部分:将n_jobs改为1或2(根据显存情况调整)
study_cat.optimize(
    objective_cat,
    n_trials=n_trials,
    n_jobs=1,
    show_progress_bar=True
)

2. 添加GPU显存限制参数

给CatBoost设置gpu_ram_part,限制单个模型使用的显存比例,避免占满显存导致后续trial无法分配资源:

params = {
    # 原有参数不变
    'task_type': 'GPU',
    'gpu_ram_part': 0.7  # 限制使用70%的GPU显存
}

3. 修正HyperbandPruner参数匹配

当前HyperbandPruner的max_resource=100与iterations的搜索范围(100-2000)不匹配,剪枝逻辑会出现异常。调整参数使其与迭代次数范围对齐:

pruner=optuna.pruners.HyperbandPruner(
    min_resource=100,  # 对应iterations下限
    max_resource=2000,  # 对应iterations上限
    reduction_factor=3
),

4. 集成CatBoost剪枝回调

使用Optuna官方的CatBoost剪枝回调,优化资源使用并确保进程正常释放GPU资源:

from optuna.integration import CatBoostPruningCallback

def objective_cat(trial):
    # 原有参数定义不变
    cat_model = CatBoostRegressor(**params, cat_features=cat_features)
    cv = KFold(n_splits=10, shuffle=True, random_state=RANDOM_SEED)
    
    # 使用剪枝回调替代原cross_val_score
    scores = []
    for fold, (train_idx, val_idx) in enumerate(cv.split(X_train, y_train)):
        X_fold_train, X_fold_val = X_train.iloc[train_idx], X_train.iloc[val_idx]
        y_fold_train, y_fold_val = y_train.iloc[train_idx], y_train.iloc[val_idx]
        
        cat_model.fit(
            X_fold_train, y_fold_train,
            eval_set=[(X_fold_val, y_fold_val)],
            callbacks=[CatBoostPruningCallback(trial, "RMSE")],
            verbose=False
        )
        scores.append(cat_model.get_evals_result()['validation_0']['RMSE'][-1])
    
    return np.mean(scores)

核心原因

错误本质是GPU显存资源耗尽:Optuna多进程并行启动多个CatBoost GPU训练任务,每个任务都需要占用显存,RTX3060的显存无法支撑多进程同时运行,导致操作系统强制终止进程。单独运行CatBoost GPU时只有一个进程,显存占用在合理范围内,因此无问题。

内容的提问来源于stack exchange,提问作者Anastasia Volkova

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 19:47:03