You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言pROC库无法生成正确多分类ROC曲线求助

Troubleshooting Multi-Class ROC/AUC Issues with Your Stacking Model

Hey there, let's figure out why your ROC curves look off and your AUC is hovering around 0.5 (which is basically random guess performance). Your stacking model has 77% accuracy, so the issue is definitely in how you're calculating and visualizing the multi-class ROC/AUC, not the model itself.

Let's Break Down the Problems

1. You're Using Hard Classifications Instead of Probabilities (Method 2)

ROC curves rely on predicted probabilities to measure how well the model distinguishes between classes. In your second method, you're using:

predictions<-as.numeric(predict(modelStack, testPredLevelOne))

This gives you hard class labels (like 1,2,3) instead of continuous probability scores. Without probability values, the model can't show gradient separation between classes—hence the flat, random-looking ROC curves.

2. Multi-Class AUC Calculation is Misaligned (Method 1)

Your first approach averages AUC values from multiclass.roc, but:

  • You might be mismatching the order of predicted probability columns with your true label classes.
  • multiclass.roc handles multi-class ROC in specific ways (One-vs-One or One-vs-Rest), and directly averaging individual AUCs without verifying class alignment leads to incorrect results.
  • If your true labels are strings (spam, not spam, undefinable) but the model's probability columns are ordered differently, you're calculating AUC for the wrong class pairs.

3. Label Encoding Inconsistencies

If your true labels are stored as unordered strings, multiclass.roc might not map them correctly to the predicted probability columns. This misalignment makes the AUC calculation meaningless.


Fixes & Corrected Code

Step 1: Standardize Label Encoding

First, ensure your true labels are factors with a consistent order that matches the model's predicted probability columns:

# Convert true labels to a factor with explicit order
testPredLevelOne$spam <- factor(testPredLevelOne$spam, 
                                levels = c("not spam", "spam", "undefinable"))

# Verify the model's probability column order (critical for alignment!)
prob_cols <- colnames(predict(modelStack, testPredLevelOne, type='prob'))
cat("Model's probability columns:", prob_cols, "\n")

Step 2: Use One-vs-Rest ROC/AUC (Correct Approach for Multi-Class)

For multi-class problems, the most intuitive method is One-vs-Rest: treat each class as the "positive" class and all others as "negative", then calculate ROC/AUC for each pair. Here's the corrected code:

library(pROC)

# Get predicted probabilities from the stacking model
pre <- predict(modelStack, testPredLevelOne, type='prob')

# Initialize variables for AUC values and plotting
auc_values <- c()
plot_colors <- c("blue", "red", "green")

# Loop through each class to build One-vs-Rest ROC curves
for (i in seq_along(prob_cols)) {
  class <- prob_cols[i]
  # Create binary true labels: 1 if the sample is the current class, 0 otherwise
  true_binary <- ifelse(testPredLevelOne$spam == class, 1, 0)
  # Get predicted probabilities for the current class
  pred_prob <- pre[[class]]
  
  # Calculate ROC curve and AUC
  roc_obj <- roc(true_binary, pred_prob)
  auc_values <- c(auc_values, roc_obj$auc)
  
  # Plot the ROC curve
  if (i == 1) {
    plot(roc_obj, main = "One-vs-Rest ROC Curves", 
         col = plot_colors[i], lwd = 2,
         xlab = "False Positive Rate", ylab = "True Positive Rate")
  } else {
    lines(roc_obj, col = plot_colors[i], lwd = 2)
  }
}

# Add legend to identify each class
legend("bottomright", legend = prob_cols, 
       col = plot_colors, lwd = 2)

# Calculate and print mean AUC (or use weighted average if class sizes are imbalanced)
mean_auc <- mean(auc_values)
cat("Mean One-vs-Rest AUC:", round(mean_auc, 4), "\n")

Step 3: Verify Stacking Data Alignment

Double-check that your level-one training data (predDF) and test data (testPredLevelOne) have the same label classes and feature columns. If testPredLevelOne has missing classes or different column orders, it can throw off the model's predictions.


Why This Works

  • By using predicted probabilities instead of hard labels, we get the continuous scores needed to draw meaningful ROC curves.
  • One-vs-Rest ensures we're evaluating the model's ability to distinguish each class from all others, which is a standard approach for multi-class ROC.
  • Standardizing label order eliminates misalignment between true labels and predicted probabilities.

内容的提问来源于stack exchange,提问作者AdeeThyag

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:34:15