如何使用自定义getROC_AUC函数结合ggplot绘制ROC曲线并实现多线展示?
Hey there! Let's walk through how to plot multiple ROC curves using your custom getROC_AUC function and ggplot2. Here's a step-by-step solution tailored to your needs:
1. Confirm Your Function & Prepare Sample Data
First, let's include your function (so we can work with it directly) and generate sample data for 3 different models—this lets you test the code immediately with reproducible results:
# Your custom ROC/AUC calculation function getROC_AUC <- function(probs, true_Y){ probsSort = sort(probs, decreasing = TRUE, index.return = TRUE) val = unlist(probsSort$x) idx = unlist(probsSort$ix) roc_y = true_Y[idx]; stack_x = cumsum(roc_y == 0)/sum(roc_y == 0) # False Positive Rate (FPR) stack_y = cumsum(roc_y == 1)/sum(roc_y == 1) # True Positive Rate (TPR) auc = sum((stack_x[2:length(roc_y)]-stack_x[1:length(roc_y)-1])*stack_y[2:length(roc_y)]) return(list(stack_x=stack_x, stack_y=stack_y, auc=auc)) } # Simulate sample data (replace this with your actual model predictions!) set.seed(123) # Ensure results are reproducible true_Y <- sample(c(0,1), size = 100, replace = TRUE) # Predictions from 3 hypothetical models prob_model1 <- runif(100, 0, 1) # Random guess-like performance prob_model2 <- ifelse(true_Y == 1, runif(50, 0.4, 1), runif(50, 0, 0.6)) # Moderate performance prob_model3 <- ifelse(true_Y == 1, runif(50, 0.6, 1), runif(50, 0, 0.4)) # Strong performance
2. Calculate ROC Metrics for Each Model
Run your function on each model's predictions to generate FPR, TPR, and AUC values:
# Compute ROC data for each model roc_model1 <- getROC_AUC(prob_model1, true_Y) roc_model2 <- getROC_AUC(prob_model2, true_Y) roc_model3 <- getROC_AUC(prob_model3, true_Y)
3. Format Data for ggplot2
ggplot works best with tidy data (one row per observation). We'll combine all ROC data into a single data frame, adding a Model column to distinguish curves—we'll even include the AUC in the model name for quick performance comparison:
# Combine into a tidy data frame roc_tidy_data <- rbind( data.frame( FPR = roc_model1$stack_x, TPR = roc_model1$stack_y, Model = paste0("Model 1 (AUC = ", round(roc_model1$auc, 3), ")") ), data.frame( FPR = roc_model2$stack_x, TPR = roc_model2$stack_y, Model = paste0("Model 2 (AUC = ", round(roc_model2$auc, 3), ")") ), data.frame( FPR = roc_model3$stack_x, TPR = roc_model3$stack_y, Model = paste0("Model 3 (AUC = ", round(roc_model3$auc, 3), ")") ) )
4. Plot Multiple ROC Curves
Now we can plot all curves in one visualization, plus add a dashed gray line to represent a random guess (AUC = 0.5) for baseline comparison:
library(ggplot2) ggplot(roc_tidy_data, aes(x = FPR, y = TPR, color = Model)) + # Draw ROC curves geom_line(linewidth = 1) + # Add random guess baseline geom_abline(intercept = 0, slope = 1, linetype = "dashed", color = "gray50") + # Customize labels and theme labs( x = "False Positive Rate (FPR)", y = "True Positive Rate (TPR)", title = "ROC Curves: Comparing Multiple Models", color = "Model" ) + theme_minimal() + theme( plot.title = element_text(hjust = 0.5, size = 14, face = "bold"), axis.title = element_text(size = 12), legend.position = "bottom" )
Quick Tips:
- Swap out the sample
prob_model1/2/3data with your actual model predictions. - The
Modelcolumn is essential—it tells ggplot which points belong to which curve. - Adjust
linewidth, colors, or theme elements to match your preferred styling.
内容的提问来源于stack exchange,提问作者Basel.D

