10个矩阵的均值与10折交叉验证下ROC绘图的预测均值需求
Alright, let's break this down step by step—you're working with 10-fold cross-validation and want to create a mean ROC curve by averaging results across all folds, right? That's a great way to get a more robust view of your model's performance instead of just looking at noisy per-fold curves. Let's walk through how to do this properly.
First, a quick note on the snippet you shared: it looks like you have single pairs of values (columns N and R, which I assume are True Positive Rate/TPR and False Positive Rate/FPR for specific thresholds) across multiple entries. To build a full mean ROC curve, you need the complete set of TPR-FPR pairs for every threshold tested in each fold—not just isolated points. If you only have these single points, you'll need to re-run your model to extract the full ROC data for each fold.
Assuming you end up with a list of ROC results (one per fold), let's use R and the pROC package (the go-to tool for ROC analysis) to handle the aggregation.
Each fold will likely have different numbers of threshold points, so we first interpolate all TPR values to a common set of FPR points (we'll use a sequence from 0 to 1 in 0.01 increments to keep things smooth). This ensures we can compute meaningful averages across folds.
Here's the code to do this:
# Install/load required packages if you haven't already install.packages(c("pROC", "dplyr", "tidyr", "ggplot2")) library(pROC) library(dplyr) library(tidyr) library(ggplot2) # Replace this with your actual 10-fold ROC data: # Each element in roc_list should be a `roc` object from pROC for one fold set.seed(123) # For reproducibility in this example roc_list <- lapply(1:10, function(fold) { # Simulate sample response and predictor data (swap with your model outputs) true_labels <- sample(c(0, 1), 150, replace = TRUE) predicted_scores <- rnorm(150, mean = true_labels, sd = 0.7) roc(true_labels ~ predicted_scores) }) # Create a common set of FPR points to interpolate all folds to common_fpr <- seq(0, 1, by = 0.01) # Interpolate TPR values for each fold at our common FPR points interpolated_tpr <- lapply(roc_list, function(roc_obj) { # Note: pROC uses specificity = 1 - FPR, so we flip the x-axis here approx(x = roc_obj$specificities, y = roc_obj$sensitivities, xout = 1 - common_fpr, method = "linear")$y }) # Convert to a data frame and calculate mean TPR across folds tpr_df <- as.data.frame(do.call(cbind, interpolated_tpr)) %>% rename_with(~paste0("fold_", .x), everything()) %>% mutate(common_fpr = common_fpr) %>% rowwise() %>% mutate(mean_tpr = mean(c_across(starts_with("fold_")), na.rm = TRUE)) %>% ungroup()
Now we can visualize the mean curve, and optionally overlay individual fold curves to see variability:
ggplot(tpr_df, aes(x = common_fpr, y = mean_tpr)) + # Plot individual fold curves (light gray for context) geom_line(data = tpr_df %>% pivot_longer(cols = starts_with("fold_"), names_to = "fold", values_to = "tpr"), aes(x = common_fpr, y = tpr, group = fold), color = "gray80", alpha = 0.4) + # Plot the mean curve (bold blue) geom_line(color = "#2c3e50", linewidth = 1.2) + # Add the random guess baseline (dashed red) geom_abline(slope = 1, intercept = 0, linetype = "dashed", color = "#e74c3c") + labs(x = "False Positive Rate (FPR)", y = "True Positive Rate (TPR)", title = "Mean ROC Curve (10-Fold Cross-Validation)", subtitle = "Individual fold curves shown in light gray") + theme_minimal()
If you want a numerical summary, compute the mean Area Under the Curve (AUC) across folds:
# Extract AUC for each fold fold_aucs <- sapply(roc_list, function(roc_obj) roc_obj$auc) # Compute mean and standard deviation of AUC mean_auc <- mean(fold_aucs) sd_auc <- sd(fold_aucs) cat(sprintf("Mean AUC across 10 folds: %.3f (Standard Deviation: %.3f)", mean_auc, sd_auc))
Quick Note on Your Partial Data
If you only have the single TPR-FPR points you shared, you can compute the mean of those points, but that won't give you a full ROC curve—it'll just be one averaged point. To get the full curve, make sure to extract all threshold-based TPR-FPR pairs from each fold's model predictions.
内容的提问来源于stack exchange,提问作者Fatima Mb

