You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Rapids CUML随机森林回归模型网格调参无评分且报CUDA错误

CUML随机森林回归模型GridSearchCV调参报错处理

问题情况

使用GridSearchCV对RAPIDS CUML的随机森林回归模型做参数调优时,无法完成参数组合的评分计算,抛出CUDA相关错误。但相同的调参方法在SVM回归模型上可以正常运行。

代码示例

from sklearn.model_selection import GridSearchCV
from sklearn.metrics import make_scorer, mean_squared_error
param_grid2 = {
    
    'bootstrap': [True, False],
    'max_depth': [16, 30, 50, 80, 90, 100 ],
    'max_features': ['auto', 'sqrt', 'log2'],
    'min_samples_leaf': [1, 3, 4, 5],
    'min_samples_split': [2,4, 8, 10, 12],
    'n_estimators': [100, 200, 300, 400 ,500]
    }

from cuml.ensemble import RandomForestRegressor

# 创建基础模型
rfr_cu = RandomForestRegressor()

# 实例化网格搜索模型
rfr_tune_cu = GridSearchCV(estimator = rfr_cu, param_grid = param_grid2, scoring='neg_mean_squared_error',  
                           cv = 3, n_jobs = -1, verbose = 3, error_score= 'raise',  return_train_score =True )

# 拟合数据时触发报错
rfr_tune_cu.fit(cu_X_train, cu_y_train)

报错信息(中文翻译)

RuntimeError                              Traceback (most recent call last)
Input In [80], in <cell line: 2>()
      1 # 拟合网格搜索到数据
----> 2 rfr_tune_cu.fit(cu_X_train, cu_y_train)

File /usr/local/lib/python3.9/dist-packages/sklearn/model_selection/_search.py:875, in BaseSearchCV.fit(self, X, y, groups, **fit_params)
    869     results = self._format_results(
    870         all_candidate_params, n_splits, all_out, all_more_results
    871     )
    873     return results
---> 875 self._run_search(evaluate_candidates)
    877 # 多指标评估在此处确定,因为如果scoring是可调用对象,返回类型只有调用后才知道
    878 first_test_score = all_out[0]["test_scores"]

File /usr/local/lib/python3.9/dist-packages/sklearn/model_selection/_search.py:1375, in GridSearchCV._run_search(self, evaluate_candidates)
   1373 def _run_search(self, evaluate_candidates):
   1374     """搜索param_grid中的所有候选参数"""
-> 1375     evaluate_candidates(ParameterGrid(self.param_grid))

File /usr/local/lib/python3.9/dist-packages/sklearn/model_selection/_search.py:822, in BaseSearchCV.fit.<locals>.evaluate_candidates(candidate_params, cv, more_results)
    814 if self.verbose > 0:
    815     print(
    816         "为{1}个候选参数各拟合{0}折,总计{2}次拟合".format(
    817             n_splits, n_candidates, n_candidates * n_splits
    818         )
    819     )
-> 822 out = parallel(
    823     delayed(_fit_and_score)(
    824         clone(base_estimator),
    825         X,
    826         y,
    827         train=train,
    828         test=test,
    829         parameters=parameters,
    830         split_progress=(split_idx, n_splits),
    831         candidate_progress=(cand_idx, n_candidates),
    832         **fit_and_score_kwargs,
    833     )
    834     for (cand_idx, parameters), (split_idx, (train, test)) in product(
    835         enumerate(candidate_params), enumerate(cv.split(X, y, groups))
    836     )
    837 )
    839 if len(out) < 1:
    840     raise ValueError(
    841         "未执行任何拟合。CV迭代器是否为空?是否没有候选参数?"
    842     )

File /usr/local/lib/python3.9/dist-packages/joblib/parallel.py:1056, in Parallel.__call__(self, iterable)
   1053     self._iterating = False
   1055 with self._backend.retrieval_context():
-> 1056     self.retrieve()
   1057 # 确保最后收到完成通知
   1058 elapsed_time = time.time() - self._start_time

File /usr/local/lib/python3.9/dist-packages/joblib/parallel.py:935, in Parallel.retrieve(self)
    933 try:
    934     if getattr(self._backend, 'supports_timeout', False):
-> 935         self._output.extend(job.get(timeout=self.timeout))
    936     else:
    937         self._output.extend(job.get())

File /usr/local/lib/python3.9/dist-packages/joblib/_parallel_backends.py:542, in LokyBackend.wrap_future_result(future, timeout)
    539 """包装Future.result以实现与multiprocessing.AsyncResults.get相同的行为"""
    541 try:
-> 542     return future.result(timeout=timeout)
    543 except CfTimeoutError as e:
    544     raise TimeoutError from e

File /usr/lib/python3.9/concurrent/futures/_base.py:446, in Future.result(self, timeout)
    444     raise CancelledError()
    445 elif self._state == FINISHED:
-> 446     return self.__get_result()
    447 else:
    448     raise TimeoutError()

File /usr/lib/python3.9/concurrent/futures/_base.py:391, in Future.__get_result(self)
    389 if self._exception:
    390     try:
-> 391         raise self._exception
    392     finally:
    393         # 打破与self._exception中异常的引用循环
    394         self = None

RuntimeError: 在以下位置遇到CUDA错误:文件=/project/cpp/src/decisiontree/batched-levelalgo/quantiles.cuh 行=52: 调用='cub::DeviceRadixSort::SortKeys( nullptr, temp_storage_bytes, data, sorted_column.data(), n_rows, 0, 8 * sizeof(T), stream)', 原因=cudaErrorInvalidDeviceFunction:无效的设备函数

解决建议

  • 版本兼容性检查:cudaErrorInvalidDeviceFunction通常是CUDA运行时版本与RAPIDS库编译时依赖的CUDA版本不匹配导致的,确认两者版本兼容。
  • 禁用多进程:GridSearchCV的n_jobs=-1会启用多进程,而CUDA上下文在多进程环境下容易出现资源冲突,将n_jobs改为1尝试。
  • 缩小参数网格:先减少参数组合数量,比如只调整n_estimators和max_depth,确认模型能正常运行后再逐步扩展参数范围。
  • 指定CUDA设备:在代码开头添加import cupy as cp; cp.cuda.Device(0).use(),强制所有运算使用同一个GPU设备,避免多进程下设备分配混乱。

内容的提问来源于stack exchange,提问作者MarMarhoun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 23:35:18