You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

probably::cal_plot_breaks指定概率估计的正确语法及报错排查

问题:使用probably::cal_plot_breaks时的报错与警告问题

我有一个包含预测类别概率和真实标签的数据框,使用probably::cal_plot_breaks指定类别概率时总是触发报错或警告,不确定是用法错误还是bug,以下是可复现代码:

library(tidyverse)
library(probably)
#> 
#> Attaching package: 'probably'
#> The following objects are masked from 'package:base':
#> 
#>     as.factor, as.ordered

set.seed(100)

test_df <- tibble(
  probability_x = runif(100),
  probability_y = 1-probability_x,
  Label = sample(
    c("x", "y"), 100, replace = TRUE
  ) %>% as.factor()
)

# 触发报错的代码
produces_error <- test_df %>% 
  cal_plot_breaks(
    truth = Label,
    estimate = probability_x
  )
#> Error in `purrr::map()`:
#> ℹ In index: 2.
#> Caused by error in `estimate_str[[.x]]`:
#> ! subscript out of bounds
#> Backtrace:
#>      ▆
#>   1. ├─test_df %>% cal_plot_breaks(truth = Label, estimate = probability_x)
#>   2. ├─probably::cal_plot_breaks(., truth = Label, estimate = probability_x)
#>   3. ├─probably:::cal_plot_breaks.data.frame(., truth = Label, estimate = probability_x)
#>   4. │ └─probably:::cal_plot_breaks_impl(...)
#>   5. │   ├─probably::.cal_table_breaks(...)
#>   6. │   └─probably:::.cal_table_breaks.data.frame(...)
#>   7. │     └─probably:::.cal_table_breaks_impl(...)
#>   8. │       └─probably:::truth_estimate_map(...)
#>   9. │         └─purrr::map(seq_along(truth_levels), ~sym(estimate_str[[.x]]))
#>  10. │           └─purrr:::map_("list", .x, .f, ..., .progress = .progress)
#>  11. │             ├─purrr:::with_indexed_errors(...)
#>  12. │             │ └─base::withCallingHandlers(...)
#>  13. │             ├─purrr:::call_with_cleanup(...)
#>  14. │             └─probably (local) .f(.x[[i]], ...)
#>  15. │               └─rlang::sym(estimate_str[[.x]])
#>  16. │                 └─rlang::is_symbol(x)
#>  17. └─purrr (local) `<fn>`(`<sbscOOBE>`)
#>  18.   └─cli::cli_abort(...)
#>  19.     └─rlang::abort(...)

# 触发警告的代码
produces_warning <- test_df %>% 
  cal_plot_breaks(
    truth = Label,
    estimate = starts_with("probability")
  )
#> Warning: Multiple class columns identified. Using: `probability_x`

原因分析

  • 报错原因:cal_plot_breaks针对多分类(含二分类)场景,要求传入所有类别对应的概率列。代码仅传入probability_x,但真实标签Label有两个水平(x、y),函数会尝试为每个标签水平匹配对应概率列,找不到第二列时触发下标越界错误。
  • 警告原因:用starts_with("probability")匹配到了两列概率,但函数默认仅选取第一列(probability_x),因此抛出警告提示多列被识别,但仅使用了第一列。

正确用法

要正常生成校准图,需传入所有类别概率列,有两种可行方式:

  1. 明确列出所有概率列(推荐,避免歧义):
test_df %>% 
  cal_plot_breaks(
    truth = Label,
    estimate = c(probability_x, probability_y)
  )
  1. 使用选择器匹配所有概率列:
    如果选择器匹配到的列数与标签水平数一致,且列顺序与标签水平顺序对应,可直接使用:
test_df %>% 
  cal_plot_breaks(
    truth = Label,
    estimate = starts_with("probability")
  )

该方式的警告可忽略,若需屏蔽可包裹suppressWarnings(),但更建议明确列名。

另外可提前验证标签水平顺序与概率列的对应关系:

levels(test_df$Label)
# 需与概率列后缀顺序一致,即c("x", "y")对应probability_x、probability_y

内容的提问来源于stack exchange,提问作者Alex Ondrus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 10:07:06