You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速RandomizedSearchCV运行?附数据场景与代码

优化RandomizedSearchCV在多层列名DataFrame上的运行速度

问题场景

用户拥有一个带多层元组类型列名的大型DataFrame:

JAN                 Feb            
            PRICE AMOUNT NAME   PRICE AMOUNT NAME
 2011-03-31     2      7    3       6      0    5
 2011-04-01     6      2    5       0      4    2
 2011-04-04     9      0    7       2      7    9
 2011-04-05     5      3    5       7      9    9

使用RandomForestRegressor结合RandomizedSearchCV调参时,出现元组列名的FutureWarning,且程序运行过慢,同时grid_rf和rscv的输入参数不可修改,需要可行的加速方案。

调参代码如下:

model = RandomForestRegressor() 
grid_rf = {
    'n_estimators': [500],  
    'max_depth': [5,10,15,20,25,30],  
    'min_samples_split': [2,5,10,15,20,25,30], 
    'min_samples_leaf': [1,5,10,15,20,25,30]
}
rscv = RandomizedSearchCV(estimator=model, param_distributions=grid_rf, cv=3, n_jobs=-1, verbose=2, n_iter=200)
rscv_fit = rscv.fit(x_train, y_train)
best_parameters = rscv_fit.best_params_
print(best_parameters)

收到的警告信息:

FutureWarning: Feature names only support names that are all strings. Got feature names with dtypes: ['tuple']. An error will be raised in 1.2. warnings.warn(

可行的加速方法

1. 修复元组列名(消除潜在性能损耗)

元组类型列名会触发Sklearn的警告,同时可能导致内部处理时的额外开销,先将列名转为字符串:

# 用下划线拼接多层列名为字符串
x_train.columns = ['_'.join(col) for col in x_train.columns]

消除警告的同时,避免不必要的性能损耗。

2. 缩小调参用的训练样本量

如果数据集规模极大,可以抽取部分样本用于调参,大幅降低计算量:

from sklearn.model_selection import train_test_split
# 抽取原训练集的20%用于调参,比例可根据需求调整
x_small, _, y_small, _ = train_test_split(x_train, y_train, train_size=0.2, random_state=42)
rscv_fit = rscv.fit(x_small, y_small)

注意:该方法可能轻微影响最优参数的泛化性,但提速效果显著。

3. 利用GPU加速(硬件层面优化)

如果有GPU资源,使用支持GPU加速的随机森林实现,比如RAPIDS的cuml库:

# 替换原模型为GPU版本
from cuml.ensemble import RandomForestRegressor
model = RandomForestRegressor()

GPU能大幅提升大型数据集的模型训练速度,适配原调参代码无需修改grid_rf和rscv参数。

4. 特征降维减少计算量

如果存在冗余特征,先通过特征重要性筛选核心特征:

# 训练基础模型获取特征重要性
base_model = RandomForestRegressor(n_estimators=500, random_state=42)
base_model.fit(x_train, y_train)
# 筛选重要性高于阈值的特征(阈值可自定义)
importances = base_model.feature_importances_
top_features = x_train.columns[importances > 0.01]
x_train_reduced = x_train[top_features]
# 基于降维后的特征调参
rscv_fit = rscv.fit(x_train_reduced, y_train)

减少输入特征数量后,每个模型的训练时间会显著缩短。

5. 优化系统资源分配

  • 确保n_jobs=-1能充分利用所有CPU核心,关闭其他占用大量资源的进程;
  • 若内存不足,可增加系统内存或使用分块训练(Sklearn的RandomForest暂不支持分块,需结合其他工具实现)。

内容的提问来源于stack exchange,提问作者AM27

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 00:53:12