probably::cal_plot_breaks指定概率估计的正确语法及报错排查
问题:使用
probably::cal_plot_breaks时的报错与警告问题 我有一个包含预测类别概率和真实标签的数据框,使用probably::cal_plot_breaks指定类别概率时总是触发报错或警告,不确定是用法错误还是bug,以下是可复现代码:
library(tidyverse) library(probably) #> #> Attaching package: 'probably' #> The following objects are masked from 'package:base': #> #> as.factor, as.ordered set.seed(100) test_df <- tibble( probability_x = runif(100), probability_y = 1-probability_x, Label = sample( c("x", "y"), 100, replace = TRUE ) %>% as.factor() ) # 触发报错的代码 produces_error <- test_df %>% cal_plot_breaks( truth = Label, estimate = probability_x ) #> Error in `purrr::map()`: #> ℹ In index: 2. #> Caused by error in `estimate_str[[.x]]`: #> ! subscript out of bounds #> Backtrace: #> ▆ #> 1. ├─test_df %>% cal_plot_breaks(truth = Label, estimate = probability_x) #> 2. ├─probably::cal_plot_breaks(., truth = Label, estimate = probability_x) #> 3. ├─probably:::cal_plot_breaks.data.frame(., truth = Label, estimate = probability_x) #> 4. │ └─probably:::cal_plot_breaks_impl(...) #> 5. │ ├─probably::.cal_table_breaks(...) #> 6. │ └─probably:::.cal_table_breaks.data.frame(...) #> 7. │ └─probably:::.cal_table_breaks_impl(...) #> 8. │ └─probably:::truth_estimate_map(...) #> 9. │ └─purrr::map(seq_along(truth_levels), ~sym(estimate_str[[.x]])) #> 10. │ └─purrr:::map_("list", .x, .f, ..., .progress = .progress) #> 11. │ ├─purrr:::with_indexed_errors(...) #> 12. │ │ └─base::withCallingHandlers(...) #> 13. │ ├─purrr:::call_with_cleanup(...) #> 14. │ └─probably (local) .f(.x[[i]], ...) #> 15. │ └─rlang::sym(estimate_str[[.x]]) #> 16. │ └─rlang::is_symbol(x) #> 17. └─purrr (local) `<fn>`(`<sbscOOBE>`) #> 18. └─cli::cli_abort(...) #> 19. └─rlang::abort(...) # 触发警告的代码 produces_warning <- test_df %>% cal_plot_breaks( truth = Label, estimate = starts_with("probability") ) #> Warning: Multiple class columns identified. Using: `probability_x`
原因分析
- 报错原因:
cal_plot_breaks针对多分类(含二分类)场景,要求传入所有类别对应的概率列。代码仅传入probability_x,但真实标签Label有两个水平(x、y),函数会尝试为每个标签水平匹配对应概率列,找不到第二列时触发下标越界错误。 - 警告原因:用
starts_with("probability")匹配到了两列概率,但函数默认仅选取第一列(probability_x),因此抛出警告提示多列被识别,但仅使用了第一列。
正确用法
要正常生成校准图,需传入所有类别概率列,有两种可行方式:
- 明确列出所有概率列(推荐,避免歧义):
test_df %>% cal_plot_breaks( truth = Label, estimate = c(probability_x, probability_y) )
- 使用选择器匹配所有概率列:
如果选择器匹配到的列数与标签水平数一致,且列顺序与标签水平顺序对应,可直接使用:
test_df %>% cal_plot_breaks( truth = Label, estimate = starts_with("probability") )
该方式的警告可忽略,若需屏蔽可包裹suppressWarnings(),但更建议明确列名。
另外可提前验证标签水平顺序与概率列的对应关系:
levels(test_df$Label) # 需与概率列后缀顺序一致,即c("x", "y")对应probability_x、probability_y
内容的提问来源于stack exchange,提问作者Alex Ondrus
相关产品推荐
相关产品推荐

