You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GridSearch中TerminatedWorkerError问题排查求助

XGBClassifier/sklearn GradientBoostingClassifier GridSearch触发TerminatedWorkerError问题

运行XGBClassifier或sklearn GradientBoostingClassifier的GridSearch网格搜索时,启动约2分钟后进程被系统终止,抛出TerminatedWorkerError。当前内存剩余约60%,但Logistic Regression、LightGBM、CatBoost及Random Forest均可正常运行。

代码示例

# XGB
xgb = XGBClassifier()
parameters = {'learning_rate': [1e-5, 1e-4, 1e-3, 1e-2, 1e-1, 1],
              'subsample'    : [0.9, 0.1],
              'n_estimators' : [20, 200],
              'max_depth'    : [2,30]
             }

grid_xgb = GridSearchCV(estimator=xgb, param_grid = parameters, cv = 5, n_jobs=-2, scoring='roc_auc')
xgb_grid = grid_xgb.fit(X_train, y_train)
y_pred=xgb_grid.predict(X_test);
print('roc_auc train:', xgb_grid.best_score_)
print('roc_auc test:', roc_auc_score(y_pred, y_test))
print("\n The best parameters across ALL searched params:\n", grid_xgb.best_params_)

错误信息

TerminatedWorkerError: A worker process managed by the executor was unexpectedly
terminated. This could be caused by a segmentation fault while calling the function
or by an excessive memory usage causing the Operating System to kill the worker.

完整警告堆栈

exception calling callback for <Future at 0x1cf50891f90 state=finished raised TerminatedWorkerError>
    Traceback (most recent call last):
      File "C:\Users\...\anaconda3\Lib\site-packages\joblib\externals\loky\_base.py", line 26, in _invoke_callbacks
        callback(self)
      File "C:\Users\...\anaconda3\Lib\site-packages\joblib\parallel.py", line 385, in __call__
        self.parallel.dispatch_next()
      File "C:\Users\...\anaconda3\Lib\site-packages\joblib\parallel.py", line 834, in dispatch_next
        if not self.dispatch_one_batch(self._original_iterator):
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "C:\Users\...\anaconda3\Lib\site-packages\joblib\parallel.py", line 901, in dispatch_one_batch
        self._dispatch(tasks)
      File "C:\Users\...\anaconda3\Lib\site-packages\joblib\parallel.py", line 819, in _dispatch
        job = self._backend.apply_async(batch, callback=cb)
              ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "C:\Users\...\anaconda3\Lib\site-packages\joblib\_parallel_backends.py", line 556, in apply_async
        future = self._workers.submit(SafeFunction(func))
                 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "C:\Users\...\anaconda3\Lib\site-packages\joblib\externals\loky\reusable_executor.py", line 176, in submit
        return super().submit(fn, *args, **kwargs)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "C:\Users\...\anaconda3\Lib\site-packages\joblib\externals\loky\process_executor.py", line 1129, in submit
        raise self._flags.broken
    joblib.externals.loky.process_executor.TerminatedWorkerError: A worker process managed by the executor was unexpectedly terminated. This could be caused by a segmentation fault while calling the function or by an excessive memory usage causing the Operating System to kill the worker.

已尝试的无效操作

  • 重装Anaconda
  • 重装XGBoost(Notebook和PowerShell中均操作过)
  • 安装MS Visual C++ 2015-2022
  • 每次操作后重启系统

使用环境

Windows 10、Intel 12400、32GB RAM


可行解决方向

  1. 限制并行线程数
    将n_jobs=-2改为更小的固定值(如n_jobs=2或n_jobs=4),避免Windows下多进程与XGBoost默认多核设置冲突。

  2. 显式指定XGBoost线程数
    初始化XGBClassifier时添加nthread参数,例如xgb = XGBClassifier(nthread=2),防止抢占GridSearch的进程资源。

  3. 缩小参数网格规模
    当前参数组合共48种,搭配5折交叉验证需执行240次训练。可先缩减参数范围(如learning_rate只保留[1e-3,1e-2,1e-1]),减少训练量排查问题。

  4. 切换多进程后端
    在代码开头添加以下代码,替换默认的loky后端:

    import joblib
    joblib.parallel_backend('multiprocessing')
    
  5. 检查数据兼容性
    确认X_train无异常值、缺失值,且数据类型均为XGBoost兼容格式(如无未处理的object类型特征),可先用小批量数据测试是否仍报错。


内容的提问来源于stack exchange,提问作者net_95

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 07:27:35