You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python的Random Forest Regressor模型代码改写为R语言?

Python随机森林回归+随机搜索调参的R语言等价实现

核心对应关系

  • Python的RandomForestRegressor → R中randomForest包的随机森林回归实现,通过caret包的train函数调用
  • Python的param_grid → R中用expand.grid生成参数组合,参数名映射:
    • n_estimators → ntree
    • max_depth → maxdepth
    • bootstrap 参数名一致
  • Python的RandomizedSearchCV → R中通过caret包的train函数结合参数抽样,实现从给定网格中随机选5组参数的逻辑

完整R语言代码

# 加载依赖包
library(randomForest)
library(caret)
library(doParallel)

# 启动并行计算(对应Python的n_jobs=-1,利用全部核心)
cl <- makePSOCKcluster(detectCores())
registerDoParallel(cl)

# 定义参数网格(完全匹配Python的param_grid)
param_grid <- expand.grid(
  mtry = floor(sqrt(ncol(x_train))),  # 随机森林默认mtry值(特征数平方根)
  ntree = c(200, 400, 600),
  maxdepth = c(10, 30, 50),
  bootstrap = c(TRUE)
)

# 从参数网格中随机抽取5组参数(对应Python的n_iter=5)
set.seed(42)
sampled_params <- param_grid[sample(nrow(param_grid), 5), ]

# 设置交叉验证与训练控制(对应Python的cv=5、verbose=2、random_state=42)
train_control <- trainControl(
  method = "cv",
  number = 5,
  verboseIter = TRUE,
  randomState = 42
)

# 训练模型(对应Python的CV.fit(x_train, y_train))
cv_model <- train(
  x = x_train,
  y = y_train,
  method = "rf",
  trControl = train_control,
  tuneGrid = sampled_params,
  verbose = TRUE
)

# 输出最佳结果(对应Python的print部分)
cat("best model:\n")
print(cv_model$bestTune)
# 回归任务默认用RMSE评估,若需R²可替换为cv_model$results$Rsquared
cat(sprintf("\nbest score: %.2f\n", max(cv_model$results$RMSE)))

# 关闭并行集群
stopCluster(cl)

补充说明

  • x_train需为数据框或矩阵格式,y_train需为数值型向量,确保数据格式符合R建模要求
  • 若无需严格匹配原参数网格的抽样逻辑,也可在trainControl中设置search="random"并指定tuneLength=5,让caret自动生成5组参数,但手动抽样方式更贴合原Python代码的参数选择逻辑

内容的提问来源于stack exchange,提问作者user19463887

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 03:10:28