You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

10个矩阵的均值与10折交叉验证下ROC绘图的预测均值需求

Alright, let's break this down step by step—you're working with 10-fold cross-validation and want to create a mean ROC curve by averaging results across all folds, right? That's a great way to get a more robust view of your model's performance instead of just looking at noisy per-fold curves. Let's walk through how to do this properly.

Step 1: Make Sure You Have Full Per-Fold ROC Data

First, a quick note on the snippet you shared: it looks like you have single pairs of values (columns N and R, which I assume are True Positive Rate/TPR and False Positive Rate/FPR for specific thresholds) across multiple entries. To build a full mean ROC curve, you need the complete set of TPR-FPR pairs for every threshold tested in each fold—not just isolated points. If you only have these single points, you'll need to re-run your model to extract the full ROC data for each fold.

Assuming you end up with a list of ROC results (one per fold), let's use R and the pROC package (the go-to tool for ROC analysis) to handle the aggregation.

Step 2: Align Thresholds Across Folds

Each fold will likely have different numbers of threshold points, so we first interpolate all TPR values to a common set of FPR points (we'll use a sequence from 0 to 1 in 0.01 increments to keep things smooth). This ensures we can compute meaningful averages across folds.

Here's the code to do this:

# Install/load required packages if you haven't already
install.packages(c("pROC", "dplyr", "tidyr", "ggplot2"))
library(pROC)
library(dplyr)
library(tidyr)
library(ggplot2)

# Replace this with your actual 10-fold ROC data:
# Each element in roc_list should be a `roc` object from pROC for one fold
set.seed(123) # For reproducibility in this example
roc_list <- lapply(1:10, function(fold) {
  # Simulate sample response and predictor data (swap with your model outputs)
  true_labels <- sample(c(0, 1), 150, replace = TRUE)
  predicted_scores <- rnorm(150, mean = true_labels, sd = 0.7)
  roc(true_labels ~ predicted_scores)
})

# Create a common set of FPR points to interpolate all folds to
common_fpr <- seq(0, 1, by = 0.01)

# Interpolate TPR values for each fold at our common FPR points
interpolated_tpr <- lapply(roc_list, function(roc_obj) {
  # Note: pROC uses specificity = 1 - FPR, so we flip the x-axis here
  approx(x = roc_obj$specificities, y = roc_obj$sensitivities,
         xout = 1 - common_fpr, method = "linear")$y
})

# Convert to a data frame and calculate mean TPR across folds
tpr_df <- as.data.frame(do.call(cbind, interpolated_tpr)) %>%
  rename_with(~paste0("fold_", .x), everything()) %>%
  mutate(common_fpr = common_fpr) %>%
  rowwise() %>%
  mutate(mean_tpr = mean(c_across(starts_with("fold_")), na.rm = TRUE)) %>%
  ungroup()
Step 3: Plot the Mean ROC Curve

Now we can visualize the mean curve, and optionally overlay individual fold curves to see variability:

ggplot(tpr_df, aes(x = common_fpr, y = mean_tpr)) +
  # Plot individual fold curves (light gray for context)
  geom_line(data = tpr_df %>% pivot_longer(cols = starts_with("fold_"), names_to = "fold", values_to = "tpr"),
            aes(x = common_fpr, y = tpr, group = fold), color = "gray80", alpha = 0.4) +
  # Plot the mean curve (bold blue)
  geom_line(color = "#2c3e50", linewidth = 1.2) +
  # Add the random guess baseline (dashed red)
  geom_abline(slope = 1, intercept = 0, linetype = "dashed", color = "#e74c3c") +
  labs(x = "False Positive Rate (FPR)", y = "True Positive Rate (TPR)",
       title = "Mean ROC Curve (10-Fold Cross-Validation)",
       subtitle = "Individual fold curves shown in light gray") +
  theme_minimal()
Step 4: Calculate Mean AUC (Optional)

If you want a numerical summary, compute the mean Area Under the Curve (AUC) across folds:

# Extract AUC for each fold
fold_aucs <- sapply(roc_list, function(roc_obj) roc_obj$auc)

# Compute mean and standard deviation of AUC
mean_auc <- mean(fold_aucs)
sd_auc <- sd(fold_aucs)

cat(sprintf("Mean AUC across 10 folds: %.3f (Standard Deviation: %.3f)", mean_auc, sd_auc))

Quick Note on Your Partial Data

If you only have the single TPR-FPR points you shared, you can compute the mean of those points, but that won't give you a full ROC curve—it'll just be one averaged point. To get the full curve, make sure to extract all threshold-based TPR-FPR pairs from each fold's model predictions.


内容的提问来源于stack exchange,提问作者Fatima Mb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:34:20