RStudio中空逻辑回归模型混淆矩阵计算异常问题咨询
Let's break down exactly what's happening here:
1. Your null model's predictions are a constant value
Your nullModel only fits an intercept term (~1), which in logistic regression estimates the log-odds of bigFire=1 based only on the proportion of positive cases in your training set. When you convert that log-odds to a probability (using type="response"), you get exactly mean(train$bigFire) — which in your case is 0.542. That's why every single test sample has the same predicted probability: the model has no features to distinguish between samples, so it just predicts the average positive rate from training data for everyone.
2. Threshold adjustments only change results if they cross that constant value
Since all predictions are identical, adjusting the threshold only has two possible outcomes:
- If your threshold is less than or equal to 0.542: Every prediction will be
TRUE(because 0.542 > threshold), so your confusion matrix will show all samples predicted as positive (like your current output where theTRUEcolumn has all 58+46=104 cases, and theFALSEcolumn is all 0). - If your threshold is greater than 0.542: Every prediction will be
FALSE, so your confusion matrix will flip to show all samples predicted as negative (theFALSEcolumn will have all 104 cases, and theTRUEcolumn is 0).
If you're only adjusting thresholds within one of these ranges (e.g., testing 0.4, 0.5, 0.54), you won't see any change in the confusion matrix — because all predictions still fall on the same side of the threshold.
3. How to verify this
Run this quick check to confirm:
# Check the proportion of positive cases in your training set mean(train$bigFire)
This should return exactly 0.542. Then try setting a threshold above that value, like 0.6:
CM <- table(test$bigFire, nullModel.pred > 0.6) print(CM)
You'll see the confusion matrix completely change to all FALSE predictions.
4. Next steps if you want threshold-dependent performance
The null model is just a baseline — it doesn't use any of your features (month, FFMC, temp, etc.). To get predictions that vary across samples (and thus see confusion matrix changes with different thresholds), you'll need to fit a logistic regression model that includes predictors, e.g.:
fullModel <- glm(bigFire ~ month + FFMC + temp + RH, data = train, family = binomial) fullModel.pred <- predict(fullModel, test, type = "response")
Now fullModel.pred will have different probabilities for different samples, so adjusting the threshold will change which samples are classified as positive/negative, and thus alter the confusion matrix.
内容的提问来源于stack exchange,提问作者Francesco De Santis

