使用Optuna调优GPU版CatBoost回归器时进程崩溃问题
背景信息
使用Optuna结合本地NVIDIA GeForce RTX 3060 GPU对CatBoost回归器做超参数调优,数据集规模为120k条训练样本、16个特征(含分类特征)。单独运行CatBoost GPU版本一切正常,但结合Optuna运行时出现进程终止错误;尝试CPU运行速度过慢,添加os.environ['CUDA_VISIBLE_DEVICES'] = '0'并移除n_jobs=-1后问题仍存在。
使用版本:CatBoost 1.2.5,Optuna 4.0.0
运行代码
RANDOM_SEED = 2906 n_trials = 100 def objective_cat(trial): params = {'iterations': trial.suggest_int('iterations', 100, 2000), 'learning_rate': trial.suggest_float('learning_rate', 1e-5, 1e-1, log=True), 'depth': trial.suggest_int('max_depth', 2, 16), 'bagging_temperature': trial.suggest_uniform('bagging_temperature', 0, 1), 'l2_leaf_reg': trial.suggest_float('l2_leaf_reg', 1e-5, 100, log=True), 'random_seed': RANDOM_SEED, 'loss_function': 'RMSE', 'task_type': 'GPU' } cat_model = CatBoostRegressor(**params, cat_features=cat_features) cv = KFold( n_splits=10, shuffle=True, random_state=RANDOM_SEED ) score = cross_val_score( cat_model, X_train, y_train, cv=cv, n_jobs=-1, scoring='neg_root_mean_squared_error' ) return -np.mean(score) study_cat = optuna.create_study(direction='minimize', sampler=optuna.samplers.TPESampler( seed=RANDOM_SEED, n_startup_trials=20, multivariate=True ), pruner=optuna.pruners.HyperbandPruner( min_resource=1, max_resource=100, reduction_factor=3 ), study_name=f'optuna_cat') optuna.logging.set_verbosity(optuna.logging.CRITICAL) study_cat.optimize( objective_cat, n_trials=n_trials, n_jobs=-1, show_progress_bar=True ) best_params_cat = study_cat.best_params
错误信息
TerminatedWorkerError Traceback (most recent call last) Cell In[70], line 39 26 study_cat = optuna.create_study(direction='minimize', 27 sampler=optuna.samplers.TPESampler( 28 seed=RANDOM_SEED, (...) 36 ), 37 study_name=f'optuna_cat') 38 optuna.logging.set_verbosity(optuna.logging.CRITICAL) ---> 39 study_cat.optimize( 40 objective_cat, 41 n_trials=n_trials, 42 n_jobs=-1, 43 show_progress_bar=True ... 764 return self._result 765 finally: TerminatedWorkerError: A worker process managed by the executor was unexpectedly terminated. This could be caused by a segmentation fault while calling the function or by an excessive memory usage causing the Operating System to kill the worker.
解决方案
1. 限制并行数量+关闭交叉验证并行
RTX3060显存有限(通常12G),n_jobs=-1会启动与CPU核心数一致的trial进程,每个进程都占用GPU显存,直接导致显存耗尽被系统终止;同时交叉验证的n_jobs=-1会进一步加剧资源冲突。
修改代码:
# 交叉验证部分:将n_jobs改为1 score = cross_val_score( cat_model, X_train, y_train, cv=cv, n_jobs=1, scoring='neg_root_mean_squared_error' ) # Optuna优化部分:将n_jobs改为1或2(根据显存情况调整) study_cat.optimize( objective_cat, n_trials=n_trials, n_jobs=1, show_progress_bar=True )
2. 添加GPU显存限制参数
给CatBoost设置gpu_ram_part,限制单个模型使用的显存比例,避免占满显存导致后续trial无法分配资源:
params = { # 原有参数不变 'task_type': 'GPU', 'gpu_ram_part': 0.7 # 限制使用70%的GPU显存 }
3. 修正HyperbandPruner参数匹配
当前HyperbandPruner的max_resource=100与iterations的搜索范围(100-2000)不匹配,剪枝逻辑会出现异常。调整参数使其与迭代次数范围对齐:
pruner=optuna.pruners.HyperbandPruner( min_resource=100, # 对应iterations下限 max_resource=2000, # 对应iterations上限 reduction_factor=3 ),
4. 集成CatBoost剪枝回调
使用Optuna官方的CatBoost剪枝回调,优化资源使用并确保进程正常释放GPU资源:
from optuna.integration import CatBoostPruningCallback def objective_cat(trial): # 原有参数定义不变 cat_model = CatBoostRegressor(**params, cat_features=cat_features) cv = KFold(n_splits=10, shuffle=True, random_state=RANDOM_SEED) # 使用剪枝回调替代原cross_val_score scores = [] for fold, (train_idx, val_idx) in enumerate(cv.split(X_train, y_train)): X_fold_train, X_fold_val = X_train.iloc[train_idx], X_train.iloc[val_idx] y_fold_train, y_fold_val = y_train.iloc[train_idx], y_train.iloc[val_idx] cat_model.fit( X_fold_train, y_fold_train, eval_set=[(X_fold_val, y_fold_val)], callbacks=[CatBoostPruningCallback(trial, "RMSE")], verbose=False ) scores.append(cat_model.get_evals_result()['validation_0']['RMSE'][-1]) return np.mean(scores)
核心原因
错误本质是GPU显存资源耗尽:Optuna多进程并行启动多个CatBoost GPU训练任务,每个任务都需要占用显存,RTX3060的显存无法支撑多进程同时运行,导致操作系统强制终止进程。单独运行CatBoost GPU时只有一个进程,显存占用在合理范围内,因此无问题。
内容的提问来源于stack exchange,提问作者Anastasia Volkova

