如何在workflow_set中使用dplyr风格选择器高效指定变量?
问题
我希望用workflow_set遍历不同变量组合构建模型。官方文档提到预处理器有三类:标准R公式、预处理前的recipe对象、dplyr风格选择器,但只给了R公式的示例。我倾向用dplyr风格,方便后续自定义函数指定结果和预测变量,但没找到相关示例。目前用recipe %>% update_role %>% update_role的方式可行但太繁琐(示例2);尝试add_variables的方式失败(示例3)。请问有没有更简洁的dplyr风格写法用于workflow_set?
示例1 R公式风格(官方示例)
set.seed(123) data <- data.frame( x1 = rnorm(100), x2 = rnorm(100), x3 = rnorm(100), x4 = rnorm(100), y = rnorm(100) ) variables<-list( first = y ~ x1, second = y ~ x2, third = y ~ x1+x2, fourth = y ~ x3+x4 ) lm_model <- linear_reg() %>% set_engine("lm") location_models <- workflow_set(preproc = variables, models = list(lm = lm_model)) location_models location_models$fit[[4]] extract_workflow(location_models, id = "third_lm") location_models <- location_models %>% mutate(fit = map(info, \(x) fit(x$workflow[[1]], data))) location_models$fit[[4]]
示例2 使用update_role的写法(可行但繁琐)
variables<-list( fist = recipe(data)%>%update_role( y, new_role = "outcome")%>%update_role( x1, new_role = "predictor"), second = recipe(data)%>%update_role( y, new_role = "outcome")%>%update_role( x2, new_role = "predictor"), third = recipe(data)%>%update_role( y, new_role = "outcome")%>%update_role( c(x1,x2), new_role = "predictor"), forth = recipe(data)%>%update_role( y, new_role = "outcome")%>%update_role( c(x3,x4), new_role = "predictor") )
示例3 尝试add_variables(失败)
variables<-list( fist = add_variables(outcomes = y, predictors = x1), second = add_variables(outcomes = y, predictors = x2), third = add_variables(outcomes = y, predictors = c(x1,x2)), forth = add_variables(outcomes = y, predictors = c(x3,x4)) )
更优的dplyr风格写法
方法1:封装自定义函数生成Recipe
通过封装函数减少重复代码,同时支持dplyr的选择语法(比如starts_with("x")这类批量选择器):
# 定义生成指定变量角色的Recipe函数 create_var_recipe <- function(outcome_col, predictor_cols) { recipe(data) %>% update_role({{outcome_col}}, new_role = "outcome") %>% update_role({{predictor_cols}}, new_role = "predictor") } # 用dplyr风格定义变量组合 variables <- list( first = create_var_recipe(y, x1), second = create_var_recipe(y, x2), third = create_var_recipe(y, c(x1, x2)), fourth = create_var_recipe(y, starts_with("x3")) # 示例:用dplyr批量选择器 ) # 构建workflow_set(和示例1逻辑一致) lm_model <- linear_reg() %>% set_engine("lm") location_models <- workflow_set(preproc = variables, models = list(lm = lm_model))
方法2:直接用Workflow结合add_variables
add_variables是workflow的步骤函数,需要包裹在workflow()中才能作为预处理器传入workflow_set:
# 用workflow封装add_variables步骤 variables <- list( first = workflow() %>% add_variables(outcomes = y, predictors = x1), second = workflow() %>% add_variables(outcomes = y, predictors = x2), third = workflow() %>% add_variables(outcomes = y, predictors = c(x1, x2)), fourth = workflow() %>% add_variables(outcomes = y, predictors = c(x3, x4)) ) # 构建workflow_set lm_model <- linear_reg() %>% set_engine("lm") location_models <- workflow_set(preproc = variables, models = list(lm = lm_model))
两种方法都能实现简洁的dplyr风格变量选择,其中自定义函数的方式更适合后续扩展预处理逻辑(比如添加特征工程步骤)。
内容的提问来源于stack exchange,提问作者Niall Marsay
相关产品推荐
相关产品推荐

