使用R语言pROC库无法生成正确多分类ROC曲线求助
Hey there, let's figure out why your ROC curves look off and your AUC is hovering around 0.5 (which is basically random guess performance). Your stacking model has 77% accuracy, so the issue is definitely in how you're calculating and visualizing the multi-class ROC/AUC, not the model itself.
Let's Break Down the Problems
1. You're Using Hard Classifications Instead of Probabilities (Method 2)
ROC curves rely on predicted probabilities to measure how well the model distinguishes between classes. In your second method, you're using:
predictions<-as.numeric(predict(modelStack, testPredLevelOne))
This gives you hard class labels (like 1,2,3) instead of continuous probability scores. Without probability values, the model can't show gradient separation between classes—hence the flat, random-looking ROC curves.
2. Multi-Class AUC Calculation is Misaligned (Method 1)
Your first approach averages AUC values from multiclass.roc, but:
- You might be mismatching the order of predicted probability columns with your true label classes.
multiclass.rochandles multi-class ROC in specific ways (One-vs-One or One-vs-Rest), and directly averaging individual AUCs without verifying class alignment leads to incorrect results.- If your true labels are strings (
spam,not spam,undefinable) but the model's probability columns are ordered differently, you're calculating AUC for the wrong class pairs.
3. Label Encoding Inconsistencies
If your true labels are stored as unordered strings, multiclass.roc might not map them correctly to the predicted probability columns. This misalignment makes the AUC calculation meaningless.
Fixes & Corrected Code
Step 1: Standardize Label Encoding
First, ensure your true labels are factors with a consistent order that matches the model's predicted probability columns:
# Convert true labels to a factor with explicit order testPredLevelOne$spam <- factor(testPredLevelOne$spam, levels = c("not spam", "spam", "undefinable")) # Verify the model's probability column order (critical for alignment!) prob_cols <- colnames(predict(modelStack, testPredLevelOne, type='prob')) cat("Model's probability columns:", prob_cols, "\n")
Step 2: Use One-vs-Rest ROC/AUC (Correct Approach for Multi-Class)
For multi-class problems, the most intuitive method is One-vs-Rest: treat each class as the "positive" class and all others as "negative", then calculate ROC/AUC for each pair. Here's the corrected code:
library(pROC) # Get predicted probabilities from the stacking model pre <- predict(modelStack, testPredLevelOne, type='prob') # Initialize variables for AUC values and plotting auc_values <- c() plot_colors <- c("blue", "red", "green") # Loop through each class to build One-vs-Rest ROC curves for (i in seq_along(prob_cols)) { class <- prob_cols[i] # Create binary true labels: 1 if the sample is the current class, 0 otherwise true_binary <- ifelse(testPredLevelOne$spam == class, 1, 0) # Get predicted probabilities for the current class pred_prob <- pre[[class]] # Calculate ROC curve and AUC roc_obj <- roc(true_binary, pred_prob) auc_values <- c(auc_values, roc_obj$auc) # Plot the ROC curve if (i == 1) { plot(roc_obj, main = "One-vs-Rest ROC Curves", col = plot_colors[i], lwd = 2, xlab = "False Positive Rate", ylab = "True Positive Rate") } else { lines(roc_obj, col = plot_colors[i], lwd = 2) } } # Add legend to identify each class legend("bottomright", legend = prob_cols, col = plot_colors, lwd = 2) # Calculate and print mean AUC (or use weighted average if class sizes are imbalanced) mean_auc <- mean(auc_values) cat("Mean One-vs-Rest AUC:", round(mean_auc, 4), "\n")
Step 3: Verify Stacking Data Alignment
Double-check that your level-one training data (predDF) and test data (testPredLevelOne) have the same label classes and feature columns. If testPredLevelOne has missing classes or different column orders, it can throw off the model's predictions.
Why This Works
- By using predicted probabilities instead of hard labels, we get the continuous scores needed to draw meaningful ROC curves.
- One-vs-Rest ensures we're evaluating the model's ability to distinguish each class from all others, which is a standard approach for multi-class ROC.
- Standardizing label order eliminates misalignment between true labels and predicted probabilities.
内容的提问来源于stack exchange,提问作者AdeeThyag

