You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Tidymodels中XGBoost的mtry按比例设置及调优报错问题

问题描述

我的训练数据原本有32列,通过recipe中的step_dummy(all_nominal_predictors(), one_hot = T)做独热编码后,建模时特征列数超过32列。此时用finalize(mtry, select(data_train, -outcome))得到的mtry上限值偏低,我希望将mtry设置为变量比例而非计数,创建自定义网格(设置比例范围)并使用tune_race_anova调优,但代码运行报错。

我的代码如下(执行recipe后):

xgb_spec <- 
  boost_tree(
             tree_depth = tune(),
             trees = 500,
             learn_rate = tune(),
             loss_reduction = tune(),
             sample_size = tune(),
             stop_iter = tune(),
             mtry = tune(),
             min_n = tune()
             ) %>% 
  set_engine("xgboost", validation = 0.2, counts = FALSE) %>% 
  set_mode("regression") 

# 通常我会有多个workflow
xgb_workflows <- 
  workflowsets::workflow_set(
    preproc = list(recipe_xg),
    models = list(xgboost = xgb_spec),
    cross = TRUE
  )

# 自定义网格
grid_xgb <- 
  grid_latin_hypercube(
    learn_rate(),
    min_n(),
    mtry(range = c(0.3, 1.0)), 
    sample_size = sample_prop(),
    tree_depth(),
    loss_reduction(),
    stop_iter(),
    size = 30
  ) 

调优代码执行时报错:

cl <- makePSOCKcluster(parallel::detectCores(logical = FALSE))
registerDoParallel(cl, cores = cl-1)

race_ctrl <-
  control_race(
     save_pred = F,
     parallel_over = "everything",
     save_workflow = F
  )

start_time <- Sys.time()

race_results <-
  xgb_workflows %>%
  workflow_map(
     "tune_race_anova",
     seed = 1503,
     resamples = train_folds,
     grid = grid_xgb, 
     control = race_ctrl
  )

end_time <- Sys.time()
training_time <- end_time-start_time

请问我遗漏了什么才能正确将mtry按比例纳入自定义网格?另外,虽然可以设置独热编码后训练数据的列数上限,但比例设置更优——因为我还有遍历不同数量预测器的场景,无需逐个更新每个workflow的mtry值。


解决方案

要让mtry按比例设置并正常运行,核心是统一参数的类型定义,具体调整如下:

1. 在模型规格中明确mtry为比例类型

boost_tree()的mtry参数默认是绝对计数类型,要切换为比例模式,需在模型规格里用mtry_prop()函数定义参数,或者在tune()中声明比例范围:

方式一:用mtry_prop()明确指定

xgb_spec <- 
  boost_tree(
             tree_depth = tune(),
             trees = 500,
             learn_rate = tune(),
             loss_reduction = tune(),
             sample_size = tune(),
             stop_iter = tune(),
             mtry = mtry_prop(range = c(0.3, 1.0)), # 明确标记为比例类型
             min_n = tune()
             ) %>% 
  set_engine("xgboost", validation = 0.2, counts = FALSE) %>% 
  set_mode("regression") 

方式二:在tune()中声明比例范围

xgb_spec <- 
  boost_tree(
             tree_depth = tune(),
             trees = 500,
             learn_rate = tune(),
             loss_reduction = tune(),
             sample_size = tune(),
             stop_iter = tune(),
             mtry = tune(range = c(0.3, 1.0)), # 直接声明比例范围
             min_n = tune()
             ) %>% 
  set_engine("xgboost", validation = 0.2, counts = FALSE) %>% 
  set_mode("regression") 

2. 调整自定义网格的mtry定义

创建网格时,要匹配模型规格的参数类型,用mtry_prop()替代mtry():

grid_xgb <- 
  grid_latin_hypercube(
    learn_rate(),
    min_n(),
    mtry_prop(range = c(0.3, 1.0)), # 与模型规格的比例类型保持一致
    sample_size = sample_prop(),
    tree_depth(),
    loss_reduction(),
    stop_iter(),
    size = 30
  ) 

3. 核心原理说明

tune包中,mtry()对应特征的绝对计数,mtry_prop()对应特征总数的比例。如果模型规格用了计数类型,网格里强行写比例范围会导致参数类型不匹配,触发报错。

统一使用mtry_prop()后,tune会自动根据每个fold中预处理后的特征总数,计算实际的mtry计数,完美适配独热编码后特征数量变化的场景,同时也能兼容不同预测器数量的workflow,无需手动更新mtry上限。

内容的提问来源于stack exchange,提问作者Andy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 09:40:35