使用tidymodels预测无标签数据:因子转NA警告处理
背景
我使用tidymodels包,按以下步骤构建模型,希望用训练好的模型对新的无标签数据(示例中的df_nolabel对象,未参与训练集或测试集)做数值预测。
无标签数据包含一个训练/测试集中不存在的因子水平。该因子变量并非模型输入变量,而是作为ID变量(recipes::update_role(Species, new_role = "ID"))用于后续可视化步骤。
问题
无标签数据中的新因子水平在预测时触发如下警告:
Warning message:
Novel level found in column "Species": "setosa".
ℹ The level has been removed, and values have been coerced to.
疑问
- 既然仍能得到预测结果,我是否可以安全忽略该警告?
- 如何避免该警告?如果可能的话,不修改workflow。
可复现代码
Sys.setenv(LANG = "en") # 必要包 library(tidyverse) library(tidymodels) # 示例数据 df = iris %>% dplyr::filter(Species != "setosa") %>% mutate(Species = factor(Species)) # 训练-测试集划分 set.seed(42) split_index <- rsample::initial_split(df, prop = 0.75, strata = "Species") train_data <- rsample::training(split_index) test_data <- rsample::testing(split_index) # 模型引擎 rf_mod <- parsnip::rand_forest( mode = "regression", engine = "ranger", # mtry 使用默认值,可查看?ranger了解默认计算方式 trees = 1000, min_n = 5) # 自定义Recipe,更新因子变量角色用于后续绘图 rf_rec <- recipes::recipe(formula = Sepal.Length ~ ., data = train_data) %>% recipes::update_role(Species, new_role = "ID") prep(rf_rec) #> #> ── Recipe ────────────────────────────────────────────────────────────────────── #> #> ── Inputs #> Number of variables by role #> outcome: 1 #> predictor: 3 #> ID: 1 #> #> ── Training information #> Training data contained 74 data points and no incomplete rows. # 工作流 rf_workflow <- workflows::workflow() %>% workflows::add_recipe(rf_rec) %>% workflows::add_model(rf_mod) # 模型性能评估,此处简化处理 model_perf <- metric_set(yardstick::rmse, yardstick::rsq) rf_last_fit <- tune::last_fit(rf_workflow, split = split_index, metrics = model_perf) # 提取训练好的模型 rf_final_model = extract_workflow(rf_last_fit) # 对无标签数据预测 df_nolabel = iris %>% dplyr::filter(Species == "setosa") %>% dplyr::select(-Sepal.Length) %>% mutate(Species = factor(Species)) predict(rf_final_model, df_nolabel) #> Warning: Novel level found in column "Species": "setosa". #> ℹ The level has been removed, and values have been coerced to <NA>. #> # A tibble: 50 × 1 #> .pred #> <dbl> #> 1 5.63 #> 2 5.58 #> 3 5.65 #> 4 5.64 #> 5 5.74 #> 6 5.74 #> 7 5.63 #> 8 5.63 #> 9 5.54 #> 10 5.64 #> # ℹ 40 more rows
Created on 2025-06-26 with reprex v2.1.1
解答
1. 能否安全忽略该警告?
可以安全忽略。因为Species被标记为ID角色,不属于模型的预测变量,它的水平变化不会影响模型的预测逻辑——模型在预测时根本不会用到这个变量。警告只是提示这个ID变量的新水平被转成了NA,但这不会干扰预测结果的生成,后续可视化时只要保留原始的Species列(不要用处理后的数据),也不会影响使用。
2. 如何避免该警告(不修改workflow)?
有两种简单可行的方法:
方法一:预测时临时移除ID列
先筛选掉Species列再预测,之后再合并回ID列,既不触发警告也不影响后续使用:# 提取模型需要的预测变量 df_pred <- df_nolabel %>% select(-Species) # 执行预测 pred_results <- predict(rf_final_model, df_pred) # 合并回ID列 final_output <- bind_cols(df_nolabel %>% select(Species), pred_results)方法二:临时设置选项抑制警告
使用rlang::with_options()强制保留ID变量的原始水平,精准抑制该类警告,不会影响其他有用警告的输出:rlang::with_options( recipes.force_terms = TRUE, predict(rf_final_model, df_nolabel) )
内容的提问来源于stack exchange,提问作者Paul

