tidymodels中step_unknown调参后未生效的技术问题
问题:step_unknown未将测试集因子变量缺失值编码为'unknown'
我使用step_unknown作为缺失值插补方法,期望将因子变量的缺失值转为新水平unknown,训练工作流后能自动处理测试集的缺失值,但实际功能未生效:
- 变量x2的缺失值未被重编码为
unknown,导致ranger随机森林模型将其强制转为NA。
构建示例数据集
set.seed(135) df <- data.frame( x1 = rnorm(100), x2 = base::sample(c('a','b','c','d'), 100, replace = T), y = base::sample(c('p','n'), size = 100, replace = T) ) na_idx <- replicate(2,sample(c(0,1),100, replace =T, prob = c(0.9,0.2))) df$x1[na_idx[,1]==1] = NA df$x2[na_idx[,2]==1] = NA colSums(is.na(df)) df_split <- initial_split(df, strata = y) df_train <- training(df_split) df_test <- testing(df_split) # 设置交叉验证折数 df_fold <- vfold_cv(df_train, v = 5)
配置含step_unknown的Recipe
base_rec <- recipe(y ~ x1 + x2, data = df) %>% step_meanimpute(all_numeric(), -all_outcomes()) %>% step_unknown(-all_outcomes(), -all_numeric())
执行prep()后查看recipe,显示已检测到x2的缺失值并标记为已处理:
> base_rec %>% prep() Data Recipe Inputs: role #variables outcome 1 predictor 2 Training data contained 100 data points and 29 incomplete rows. Operations: Mean Imputation for x1 [trained] Unknown factor level assignment for x2 [trained]
配置随机森林模型与工作流
# 设置模型引擎 rf_eng <- rand_forest( mtry = tune(), trees = 500, min_n = tune(), mode = 'classification' ) %>% set_engine('ranger', importance = 'impurity') # 构建工作流 wflow_rf <- workflow() %>% add_model(rf_eng) %>% add_recipe(base_rec)
调参执行
tune_res <- tune_grid( wflow_rf, resamples = df_fold, grid = 2, metrics = metric_set(accuracy,roc_auc,f_meas), control = control_grid(verbose = T) )
问题现象
调参结果显示第2折中x2的缺失值未被正确编码,被识别为新水平:
> tune_res$.notes [[1]] # A tibble: 0 x 1 # … with 1 variable: .notes <chr> [[2]] # A tibble: 2 x 1 .notes <chr> 1 model 1/2 (predictions): Novel levels found in column 'x2': NA. The levels have been removed, and values have been coerced to 'NA'. 2 model 2/2 (predictions): Novel levels found in column 'x2': NA. The levels have been removed, and values have been coerced to 'NA'.
查看第2折的训练集与测试集,训练集存在NA,但recipe未处理测试集的x2缺失值:
训练集x2分布
> df_fold$splits$`2`$data %>% + slice(df_fold$splits$`2`$in_id) %>% + select(x2) %>% + table(useNA = 'ifany') . a b c d <NA> 14 18 11 11 7
测试集x2分布
> df_fold$splits$`2`$data %>% + slice(-df_fold$splits$`2`$in_id) %>% + select(x2) %>% + table(useNA = 'ifany') . a b c d <NA> 3 5 2 0 5
解决方法
问题出在step_unknown的默认行为:它仅处理训练集中出现的新因子水平,不会自动识别并处理NA。需要显式设置参数,让它同时处理训练集和测试集中的NA值:
修改后的Recipe:
base_rec <- recipe(y ~ x1 + x2, data = df) %>% step_meanimpute(all_numeric(), -all_outcomes()) %>% step_unknown(-all_outcomes(), -all_numeric(), new_level = 'unknown', na = TRUE)
参数说明:
na = TRUE:指定将NA值视为需要编码为新水平的情况new_level = 'unknown':明确设置编码后的新水平名称(默认值即为'unknown',显式设置更清晰)
修改后重新运行工作流,测试集的NA值会被正确编码为unknown,不会再出现ranger识别到新水平NA的警告。
内容的提问来源于stack exchange,提问作者PPho
相关产品推荐
相关产品推荐

