在R中为WBCD数据集的GLM与LDA模型选阈值实现100%灵敏度
Great question—let's break this down step by step, since getting 100% sensitivity (no missed malignant cases) is critical for breast cancer screening, and threshold tuning is exactly the right tool for the job.
First: What's a Threshold, Anyway?
When you train a logistic regression (or LDA) model, it outputs a probability (between 0 and 1) that a sample is malignant (M). The default threshold of 0.5 means:
- If the probability > 0.5: predict
M - Else: predict
B(benign)
Sensitivity (true positive rate) measures how many actual M samples we correctly identify. To get 100% sensitivity, we need every single actual M sample to be predicted as M—which means we need to lower the threshold enough so that all actual M samples have a predicted probability above that threshold.
Why Your Current Adjustments Aren't Working
Looking at your code, there are two likely issues:
- Typos: You wrote
dim(tested)[1]but your test dataset is namedtestdf—this would throw an error, so you probably fixed that but didn't update the code snippet. - Guessing thresholds instead of targeting the right value: Just tweaking 0.5 to 0.4 or 0.3 might not hit the exact threshold needed to capture all
Msamples. We need to calculate the exact threshold based on your model's predictions.
Step-by-Step Fix for Logistic Regression
Here's how to adjust your code to guarantee 100% sensitivity:
# Clean up the dataset df <- subset(df, select = -c(X, id)) # Train-test split (fixed, same as your code) set.seed(4) nrows <- NROW(df) index <- sample(1:nrows, 0.7 * nrows) traindf <- df[index,] testdf <- df[-index,] # Fit the logistic regression model glm.fit <- glm(diagnosis~., data=traindf, family=binomial) glm.probs <- predict(glm.fit, testdf, type="response") # Find the minimum predicted probability for actual malignant samples actual_m_probs <- glm.probs[testdf$diagnosis == "M"] # Set threshold slightly below this minimum to avoid edge cases (e.g., exact equality) threshold <- min(actual_m_probs) - 1e-6 # Generate predictions with the new threshold glm.pred <- ifelse(glm.probs > threshold, "M", "B") # Print the updated confusion matrix conf_matrix <- table(glm.pred, testdf$diagnosis) print(conf_matrix)
What This Does
actual_m_probsgrabs all predicted probabilities for samples that are actuallyMin your test set. The smallest value here is the lowest probability your model assigns to a true malignant case.- By setting the threshold just below that minimum, we ensure every true
Msample will have a probability above the threshold—so they'll all be predicted asM, making the top-right value of your confusion matrix (false negatives) 0.
Tradeoff to Keep in Mind
100% sensitivity comes at a cost: you'll likely have more false positives (benign samples predicted as malignant, the bottom-left value in your confusion matrix). This is the classic accuracy-sensitivity tradeoff, but for cancer screening, prioritizing no missed cases is often worth it.
Applying This to LDA
The exact same logic works for LDA! When you fit an LDA model, use predict(lda.fit, testdf, type="response") to get probabilities, then repeat the steps above to find the threshold that captures all true M samples.
A Note on "Introduction to Statistical Learning"
You're right—Chapter 4 doesn't dive deep into threshold tuning, but Chapter 5 covers ROC curves, which visualize the tradeoff between sensitivity and specificity across all possible thresholds. You can use the pROC package to plot this and find the exact threshold for 100% sensitivity:
library(pROC) roc_obj <- roc(testdf$diagnosis, glm.probs) plot(roc_obj) # Shows sensitivity vs specificity for every threshold # Get the threshold that gives 100% sensitivity coords(roc_obj, x="sensitivity", input=1, ret="threshold")
This will give you the precise threshold value you need, just like the manual calculation above.
内容的提问来源于stack exchange,提问作者Federico Sparapan

