You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidymodels预测无标签数据:因子转NA警告处理

问题解答

背景

我使用tidymodels包,按以下步骤构建模型,希望用训练好的模型对新的无标签数据(示例中的df_nolabel对象,未参与训练集或测试集)做数值预测。

无标签数据包含一个训练/测试集中不存在的因子水平。该因子变量并非模型输入变量,而是作为ID变量(recipes::update_role(Species, new_role = "ID"))用于后续可视化步骤。

问题

无标签数据中的新因子水平在预测时触发如下警告:

Warning message:
Novel level found in column "Species": "setosa".
ℹ The level has been removed, and values have been coerced to .

疑问

  1. 既然仍能得到预测结果,我是否可以安全忽略该警告?
  2. 如何避免该警告?如果可能的话,不修改workflow。

可复现代码

Sys.setenv(LANG = "en")
# 必要包
library(tidyverse)
library(tidymodels)

# 示例数据
df = iris %>% dplyr::filter(Species != "setosa") %>% mutate(Species = factor(Species))

# 训练-测试集划分
set.seed(42)
split_index <- rsample::initial_split(df, prop = 0.75, strata = "Species")
train_data <- rsample::training(split_index)
test_data  <- rsample::testing(split_index)

# 模型引擎
rf_mod <- parsnip::rand_forest(
  mode = "regression",
  engine = "ranger",
  # mtry 使用默认值,可查看?ranger了解默认计算方式
  trees = 1000,
  min_n = 5)

# 自定义Recipe,更新因子变量角色用于后续绘图
rf_rec <- recipes::recipe(formula = Sepal.Length ~ ., data = train_data) %>%
  recipes::update_role(Species, new_role = "ID")
prep(rf_rec)
#> 
#> ── Recipe ──────────────────────────────────────────────────────────────────────
#> 
#> ── Inputs
#> Number of variables by role
#> outcome:   1
#> predictor: 3
#> ID:        1
#> 
#> ── Training information
#> Training data contained 74 data points and no incomplete rows.

# 工作流
rf_workflow <- workflows::workflow() %>% 
  workflows::add_recipe(rf_rec) %>% 
  workflows::add_model(rf_mod)

# 模型性能评估,此处简化处理
model_perf <- metric_set(yardstick::rmse, yardstick::rsq)

rf_last_fit <- tune::last_fit(rf_workflow, split = split_index, metrics = model_perf)

# 提取训练好的模型
rf_final_model = extract_workflow(rf_last_fit)

# 对无标签数据预测
df_nolabel = iris %>% 
  dplyr::filter(Species == "setosa") %>% 
  dplyr::select(-Sepal.Length) %>% 
  mutate(Species = factor(Species))

predict(rf_final_model, df_nolabel)
#> Warning: Novel level found in column "Species": "setosa".
#> ℹ The level has been removed, and values have been coerced to <NA>.
#> # A tibble: 50 × 1
#>    .pred
#>    <dbl>
#>  1  5.63
#>  2  5.58
#>  3  5.65
#>  4  5.64
#>  5  5.74
#>  6  5.74
#>  7  5.63
#>  8  5.63
#>  9  5.54
#> 10  5.64
#> # ℹ 40 more rows

Created on 2025-06-26 with reprex v2.1.1


解答

1. 能否安全忽略该警告?

可以安全忽略。因为Species被标记为ID角色,不属于模型的预测变量,它的水平变化不会影响模型的预测逻辑——模型在预测时根本不会用到这个变量。警告只是提示这个ID变量的新水平被转成了NA,但这不会干扰预测结果的生成,后续可视化时只要保留原始的Species列(不要用处理后的数据),也不会影响使用。

2. 如何避免该警告(不修改workflow)?

有两种简单可行的方法:

  • 方法一:预测时临时移除ID列
    先筛选掉Species列再预测,之后再合并回ID列,既不触发警告也不影响后续使用:

    # 提取模型需要的预测变量
    df_pred <- df_nolabel %>% select(-Species)
    # 执行预测
    pred_results <- predict(rf_final_model, df_pred)
    # 合并回ID列
    final_output <- bind_cols(df_nolabel %>% select(Species), pred_results)
    
  • 方法二:临时设置选项抑制警告
    使用rlang::with_options()强制保留ID变量的原始水平,精准抑制该类警告,不会影响其他有用警告的输出:

    rlang::with_options(
      recipes.force_terms = TRUE,
      predict(rf_final_model, df_nolabel)
    )
    

内容的提问来源于stack exchange,提问作者Paul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 19:34:59