You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用pROC阈值生成混淆矩阵,计算统计量95%置信区间?

用pROC阈值生成混淆矩阵并计算统计量95%置信区间

步骤1:提取pROC的最佳阈值

先运行ROC分析代码,通过coords()获取最佳阈值及对应统计量:

library(pROC)
library(tibble)
library(epiR)

# 示例数据
data <- tribble(
  ~death, ~score,
  0, 0.132,
  1, 0.19, 
  0, 0.03,
  1, 0.131,
  0, 0.02
)

# 拟合ROC曲线
roc <- roc(data$death, data$score, smoothed = TRUE,
           ci=TRUE, ci.alpha=0.95, stratified=FALSE,
           plot=TRUE, auc.polygon=TRUE, max.auc.polygon=TRUE, grid=TRUE,
           print.auc=TRUE, show.thres=TRUE)

# 获取最佳阈值(默认用Youden指数筛选)
best_coords <- coords(roc, x="best", ret=c("threshold", "specificity", "sensitivity"))
best_threshold <- best_coords["threshold"]

步骤2:基于阈值生成预测分类

先查看ROC的方向(避免搞反阳性/阴性判断规则),再生成预测标签:

# 查看ROC方向,明确分类逻辑
cat("ROC方向:", roc$direction, "\n")
# 示例中direction应为">",表示case组(death=1)的score更高,因此score >= 阈值时预测为1

# 生成预测标签
data$predicted <- ifelse(data$score >= best_threshold, 1, 0)

步骤3:构建混淆矩阵并计算置信区间

用真实标签和预测标签生成混淆矩阵,再通过epiR::epi.tests()计算带95%CI的统计量:

# 构建混淆矩阵:行=真实标签,列=预测标签
conf_mat <- table(真实死亡状态=data$death, 预测死亡状态=data$predicted)
print("混淆矩阵:")
print(conf_mat)

# 计算带95%置信区间的统计量
stat_result <- epi.tests(conf_mat, conf.level = 0.95)
print("带95%置信区间的统计结果:")
print(stat_result)

关键说明

  • 若coords()返回多个最佳阈值(多个阈值对应相同的最佳指标),可通过coords()的best.method参数明确筛选规则(如best.method="youden"或best.method="closest.topleft")
  • 必须匹配ROC的direction设置分类逻辑,否则会导致混淆矩阵完全错误
  • epi.tests()输出结果包含灵敏度、特异度、准确率、阳性预测值(PPV)、阴性预测值(NPV)等统计量的95%置信区间,直接满足期刊要求

内容的提问来源于stack exchange,提问作者John Ryan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 08:10:26