如何从caret十折交叉验证中提取训练与测试集AUROC
Great question—when working with tiny datasets, checking how your model performs on each fold's training data vs its validation data is one of the best ways to spot overfitting. Let's walk through exactly how to get those per-fold training AUROC values with your caret workflow.
Step 1: Recreate the exact fold splits
First, we need to replicate the 10-fold partitions that your original train() call used. This ensures we're evaluating the same training subsets as your initial cross-validation run:
# Use the SAME seed as your original code to match fold splits set.seed(998) folds <- createFolds(my_data$Class, k = 10, returnTrain = TRUE)
Step 2: Calculate training AUROC for each fold
We'll use the optimal hyperparameters your model selected (model$bestTune) to fit a model on each fold's training data, then compute the AUROC for that training set. We'll use the pROC package to calculate AUROC:
library(pROC) # Initialize a vector to store our training AUROC values train_auroc <- numeric(length(folds)) # Loop through each fold for (i in seq_along(folds)) { # Grab the training indices for this fold train_idx <- folds[[i]] train_subset <- my_data[train_idx, ] # Fit the xgbTree model with the optimal parameters from your original run fold_model <- train( Class ~ ., data = train_subset, method = "xgbTree", trControl = trainControl(method = "none"), # No extra CV here—just fit once tuneGrid = model$bestTune, metric = "ROC" ) # Get predicted probabilities for the training subset train_probs <- predict(fold_model, train_subset, type = "prob") # Calculate AUROC (adjust 'M' to match your positive class if needed) # Check levels(my_data$Class) to confirm your class labels—Sonar uses "M" and "R" train_roc <- roc(response = train_subset$Class, predictor = train_probs$M) train_auroc[i] <- auc(train_roc) } # View the per-fold training AUROCs train_auroc # Combine with validation AUROCs for easy comparison performance_comparison <- data.frame( Fold = 1:10, Training_AUROC = train_auroc, Validation_AUROC = model$resample$ROC ) print(performance_comparison)
Key things to note:
- Consistent splits: Using the same seed as your original code ensures we're working with identical training/validation partitions—this is non-negotiable for reliable comparisons.
- Optimal parameters: Reusing
model$bestTunemeans each fold's model uses the exact hyperparameters your initial CV selected. - Overfitting check: If your training AUROC is consistently near 1.0 but validation AUROC is much lower, that's a clear red flag for overfitting—especially with a small dataset like Sonar.
Alternative (trickier) method: Using savePredictions = "all"
If you want to avoid refitting models, you can modify your original trainControl to save all predictions (including training sets) by setting savePredictions = "all". However, filtering the saved predictions to isolate training set results per fold is more complex and easy to mess up. The refitting method above is more straightforward and less error-prone for your use case.
内容的提问来源于stack exchange,提问作者Keshav M

