You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

tidymodels中step_unknown调参后未生效的技术问题

问题:step_unknown未将测试集因子变量缺失值编码为'unknown'

我使用step_unknown作为缺失值插补方法,期望将因子变量的缺失值转为新水平unknown,训练工作流后能自动处理测试集的缺失值,但实际功能未生效:

  • 变量x2的缺失值未被重编码为unknown,导致ranger随机森林模型将其强制转为NA。

构建示例数据集

set.seed(135)
df <- data.frame(
  x1 = rnorm(100),
  x2 = base::sample(c('a','b','c','d'), 100, replace = T),
  y = base::sample(c('p','n'), size = 100, replace = T)
)

na_idx <- replicate(2,sample(c(0,1),100, replace =T, prob = c(0.9,0.2)))

df$x1[na_idx[,1]==1] = NA
df$x2[na_idx[,2]==1] = NA

colSums(is.na(df))

df_split <-  initial_split(df, strata = y)
df_train <- training(df_split)
df_test <- testing(df_split)
  
# 设置交叉验证折数
df_fold <- vfold_cv(df_train, v = 5)

配置含step_unknown的Recipe

base_rec <- recipe(y ~ x1 + x2,
                 data = df) %>%
  step_meanimpute(all_numeric(), -all_outcomes()) %>%
  step_unknown(-all_outcomes(), -all_numeric())

执行prep()后查看recipe,显示已检测到x2的缺失值并标记为已处理:

> base_rec %>% prep()
Data Recipe

Inputs:

      role #variables
   outcome          1
 predictor          2

Training data contained 100 data points and 29 incomplete rows. 

Operations:

Mean Imputation for x1 [trained]
Unknown factor level assignment for x2 [trained]

配置随机森林模型与工作流

# 设置模型引擎
rf_eng <- rand_forest(
  mtry = tune(),
  trees = 500,
  min_n = tune(),
  mode = 'classification'
  ) %>% 
  set_engine('ranger', importance = 'impurity')

# 构建工作流
wflow_rf <- 
  workflow() %>% 
  add_model(rf_eng) %>% 
  add_recipe(base_rec)

调参执行

tune_res <- tune_grid(
  wflow_rf,
  resamples = df_fold,
  grid = 2,
  metrics = metric_set(accuracy,roc_auc,f_meas),
  control = control_grid(verbose = T)
)

问题现象

调参结果显示第2折中x2的缺失值未被正确编码,被识别为新水平:

> tune_res$.notes
[[1]]
# A tibble: 0 x 1
# … with 1 variable: .notes <chr>

[[2]]
# A tibble: 2 x 1
  .notes                                                                                                                             
  <chr>                                                                                                                              
1 model 1/2 (predictions): Novel levels found in column 'x2': NA. The levels have been removed, and values have been coerced to 'NA'.
2 model 2/2 (predictions): Novel levels found in column 'x2': NA. The levels have been removed, and values have been coerced to 'NA'.

查看第2折的训练集与测试集,训练集存在NA,但recipe未处理测试集的x2缺失值:

训练集x2分布

> df_fold$splits$`2`$data %>% 
+     slice(df_fold$splits$`2`$in_id) %>% 
+     select(x2) %>% 
+     table(useNA = 'ifany')
.
   a    b    c    d <NA> 
  14   18   11   11    7 

测试集x2分布

> df_fold$splits$`2`$data %>% 
+     slice(-df_fold$splits$`2`$in_id) %>% 
+     select(x2) %>% 
+     table(useNA = 'ifany')
.
   a    b    c    d <NA> 
   3    5    2    0    5 

解决方法

问题出在step_unknown的默认行为:它仅处理训练集中出现的新因子水平,不会自动识别并处理NA。需要显式设置参数,让它同时处理训练集和测试集中的NA值:

修改后的Recipe:

base_rec <- recipe(y ~ x1 + x2, data = df) %>%
  step_meanimpute(all_numeric(), -all_outcomes()) %>%
  step_unknown(-all_outcomes(), -all_numeric(), 
               new_level = 'unknown', na = TRUE)

参数说明:

  • na = TRUE:指定将NA值视为需要编码为新水平的情况
  • new_level = 'unknown':明确设置编码后的新水平名称(默认值即为'unknown',显式设置更清晰)

修改后重新运行工作流,测试集的NA值会被正确编码为unknown,不会再出现ranger识别到新水平NA的警告。

内容的提问来源于stack exchange,提问作者PPho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 09:45:45