You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中为WBCD数据集的GLM与LDA模型选阈值实现100%灵敏度

Threshold Tuning for 100% Sensitivity in Wisconsin Breast Cancer Classification

Great question—let's break this down step by step, since getting 100% sensitivity (no missed malignant cases) is critical for breast cancer screening, and threshold tuning is exactly the right tool for the job.

First: What's a Threshold, Anyway?

When you train a logistic regression (or LDA) model, it outputs a probability (between 0 and 1) that a sample is malignant (M). The default threshold of 0.5 means:

  • If the probability > 0.5: predict M
  • Else: predict B (benign)

Sensitivity (true positive rate) measures how many actual M samples we correctly identify. To get 100% sensitivity, we need every single actual M sample to be predicted as M—which means we need to lower the threshold enough so that all actual M samples have a predicted probability above that threshold.

Why Your Current Adjustments Aren't Working

Looking at your code, there are two likely issues:

  1. Typos: You wrote dim(tested)[1] but your test dataset is named testdf—this would throw an error, so you probably fixed that but didn't update the code snippet.
  2. Guessing thresholds instead of targeting the right value: Just tweaking 0.5 to 0.4 or 0.3 might not hit the exact threshold needed to capture all M samples. We need to calculate the exact threshold based on your model's predictions.

Step-by-Step Fix for Logistic Regression

Here's how to adjust your code to guarantee 100% sensitivity:

# Clean up the dataset
df <- subset(df, select = -c(X, id)) 

# Train-test split (fixed, same as your code)
set.seed(4)
nrows <- NROW(df)
index <- sample(1:nrows, 0.7 * nrows)
traindf <- df[index,]
testdf <- df[-index,]

# Fit the logistic regression model
glm.fit <- glm(diagnosis~., data=traindf, family=binomial)
glm.probs <- predict(glm.fit, testdf, type="response")

# Find the minimum predicted probability for actual malignant samples
actual_m_probs <- glm.probs[testdf$diagnosis == "M"]
# Set threshold slightly below this minimum to avoid edge cases (e.g., exact equality)
threshold <- min(actual_m_probs) - 1e-6

# Generate predictions with the new threshold
glm.pred <- ifelse(glm.probs > threshold, "M", "B")

# Print the updated confusion matrix
conf_matrix <- table(glm.pred, testdf$diagnosis)
print(conf_matrix)

What This Does

  • actual_m_probs grabs all predicted probabilities for samples that are actually M in your test set. The smallest value here is the lowest probability your model assigns to a true malignant case.
  • By setting the threshold just below that minimum, we ensure every true M sample will have a probability above the threshold—so they'll all be predicted as M, making the top-right value of your confusion matrix (false negatives) 0.

Tradeoff to Keep in Mind

100% sensitivity comes at a cost: you'll likely have more false positives (benign samples predicted as malignant, the bottom-left value in your confusion matrix). This is the classic accuracy-sensitivity tradeoff, but for cancer screening, prioritizing no missed cases is often worth it.

Applying This to LDA

The exact same logic works for LDA! When you fit an LDA model, use predict(lda.fit, testdf, type="response") to get probabilities, then repeat the steps above to find the threshold that captures all true M samples.

A Note on "Introduction to Statistical Learning"

You're right—Chapter 4 doesn't dive deep into threshold tuning, but Chapter 5 covers ROC curves, which visualize the tradeoff between sensitivity and specificity across all possible thresholds. You can use the pROC package to plot this and find the exact threshold for 100% sensitivity:

library(pROC)
roc_obj <- roc(testdf$diagnosis, glm.probs)
plot(roc_obj) # Shows sensitivity vs specificity for every threshold
# Get the threshold that gives 100% sensitivity
coords(roc_obj, x="sensitivity", input=1, ret="threshold")

This will give you the precise threshold value you need, just like the manual calculation above.

内容的提问来源于stack exchange,提问作者Federico Sparapan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:01:55