使用GridSearchCV训练cuML随机森林回归时遇评分失败与类型错误
尝试用GridSearchCV训练cuML的随机森林(cuRFr)回归模型,已将所有数据类型转换为float32,但仍触发大量警告与类型错误。
代码实现
combined_df=cpd.concat([train_df,evaluate_df]) combined_df = combined_df.astype({ 'Mcap_w': 'float32', 'constant': 'int32', 'TotalAssets': 'float32', 'NItoCommon_w': 'float32', 'NIbefEIPrefDiv_w': 'float32', 'PrefDiv_w': 'float32', },error='raise') print(combined_df.iloc[:,2:].info(),combined_df['Mcap_w']) test_fold = [0] * len(train_df) + [1] * len(evaluate_df) #p_test3 = {'n_estimators':[50,100,200,300,500],'max_depth':[3,4,5,6,7,8], 'max_features':[5,10,15,21]} p_test3 = {'n_estimators':[20,50,200,500],'max_depth':[3,5,7,10], 'max_features':[25]} tuning = GridSearchCV(estimator =cuRFr(n_streams=1, min_samples_split=2, min_samples_leaf=1, random_state=0), param_grid = p_test3, scoring='r2', cv=PredefinedSplit(test_fold=test_fold)) tuning.fit(combined_df.iloc[:,2:],combined_df['Mcap_w']) print(tuning.best_score_) tuning.cv_results_, tuning.best_params_, tuning.best_score_
数据类型输出
# Column Non-Null Count Dtype --- ------ -------------- ----- 0 TotalAssets 896 non-null float32 1 NItoCommon_w 896 non-null float32 2 NIbefEIPrefDiv_w 896 non-null float32 dtypes: float32(3) memory usage: 101.5 KB None
执行print(combined_df['Mcap_w'])返回:Name: Mcap_w, Length: 896, dtype: float32
错误与警告信息
miniconda3/envs/rapid/lib/python3.10/site-packages/sklearn/model_selection/_validation.py:988: UserWarning: 评分失败。该训练-测试分区在当前参数下的得分将被设为nan。
TypeError: 不允许通过
array隐式转换为主机端NumPy数组。若要显式构建GPU矩阵,请考虑使用.to_cupy();若要显式构建主机端矩阵,请考虑使用.to_numpy()。miniconda3/envs/rapid/lib/python3.10/site-packages/sklearn/model_selection/_validation.py:988: UserWarning: 评分失败。该训练-测试分区在当前参数下的得分将被设为nan。
miniconda3/envs/rapid/lib/python3.10/site-packages/cudf/core/frame.py", line 402, in
array
raise TypeError( TypeError: 不允许通过array隐式转换为主机端NumPy数组。若要显式构建GPU矩阵,请考虑使用.to_cupy();若要显式构建主机端矩阵,请考虑使用.to_numpy()。
解决方案
核心原因
scikit-learn的GridSearchCV不原生支持cuDF的GPU数据格式,在执行评分逻辑时会尝试隐式将cuDF对象转换为NumPy数组,而cuDF默认禁止这种隐式转换,因此触发报错。
解决方法
方法一:转换为NumPy数组(CPU数据)
适合数据量较小的场景,在调用fit时显式将数据转为NumPy数组:# 修改fit行 tuning.fit(combined_df.iloc[:,2:].to_numpy(), combined_df['Mcap_w'].to_numpy())方法二:使用cuML原生GridSearchCV
cuML提供了兼容GPU数据的GridSearchCV实现,无需转换数据,还能保持GPU加速优势:# 替换导入 from cuml.model_selection import GridSearchCV # 实例化和fit保持原有逻辑即可 tuning = GridSearchCV(estimator =cuRFr(n_streams=1, min_samples_split=2, min_samples_leaf=1, random_state=0), param_grid = p_test3, scoring='r2', cv=PredefinedSplit(test_fold=test_fold)) tuning.fit(combined_df.iloc[:,2:], combined_df['Mcap_w'])
注意事项
若数据集规模较大,方法一可能导致CPU内存不足,此时优先选择方法二,充分利用GPU算力。
内容的提问来源于stack exchange,提问作者Mostafa Bouzari

