使用R语言mlr包绘制二分类目标偏依赖图遇问题求助
Fixing Partial Dependence Plots (PDPs) Showing Class Labels Instead of Probabilities in mlr
Hey there! Let's sort out why your partial dependence plots are only showing class labels instead of the probabilities you need. The issue almost certainly ties to how you've configured your classification task, learner, and partial dependence data generation—let's break down the fixes step by step.
Key Issues in Your Original Workflow
From your code snippet, two critical pieces are likely missing:
- Your learner isn't configured to output probabilities (it's defaulting to class labels).
- The partial dependence data generation isn't set up to pull probability values instead of categorical predictions.
Corrected Full Workflow
Here's the adjusted code with explanations for each fix:
library(mlr) library(dplyr) library(ranger) # Prepare binary classification dataset iris_bin <- iris %>% filter(Species != "virginica") %>% mutate(bin_target = ifelse(Species == "setosa", TRUE, FALSE)) %>% select(-Species) # 1. Define classification task with explicit positive class task_bin <- makeClassifTask( data = iris_bin, target = "bin_target", positive = "TRUE" # Critical: defines which class we want probabilities for ) # 2. Configure learner to output probabilities (not just class labels) lrn_ranger <- makeLearner( "classif.ranger", predict.type = "prob", # Core fix: tells the model to return probabilities num.trees = 100 ) # Train the model model_bin <- train(lrn_ranger, task_bin) # 3. Generate partial dependence data with probability extraction pdp_data <- generatePartialDependenceData( model = model_bin, task = task_bin, features = "Sepal.Length", # Replace with your target feature # Custom function to pull the positive class probability fun = function(model, newdata) { predict(model, newdata = newdata, type = "prob")$data[, "TRUE"] } ) # Plot the PDP—this will now show probability values! plotPartialDependence(pdp_data)
Why This Works
Let's break down the critical changes:
- Explicit Positive Class: The
positive = "TRUE"argument inmakeClassifTaskclarifies which category we're calculating probabilities for, eliminating ambiguity in binary classification. - Probability-Focused Learner: Setting
predict.type = "prob"inmakeLearnerensures the ranger model outputs continuous probability values instead of discrete class labels. - Custom Probability Extraction: The
funparameter ingeneratePartialDependenceDataexplicitly pulls the probability column for your positive class, ensuring the PDP data uses numerical values instead of categorical labels.
Quick Troubleshooting Check
If you still run into issues:
- Verify that your
bin_targetcolumn is a logical or factor (mlr handles these correctly for binary classification). - Double-check that the feature name in
features = "Sepal.Length"matches exactly with your dataset's column names.
内容的提问来源于stack exchange,提问作者stats-hb
相关产品推荐
相关产品推荐

