You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RStudio中空逻辑回归模型混淆矩阵计算异常问题咨询

Why adjusting the threshold doesn't change your null model's confusion matrix

Let's break down exactly what's happening here:

1. Your null model's predictions are a constant value

Your nullModel only fits an intercept term (~1), which in logistic regression estimates the log-odds of bigFire=1 based only on the proportion of positive cases in your training set. When you convert that log-odds to a probability (using type="response"), you get exactly mean(train$bigFire) — which in your case is 0.542. That's why every single test sample has the same predicted probability: the model has no features to distinguish between samples, so it just predicts the average positive rate from training data for everyone.

2. Threshold adjustments only change results if they cross that constant value

Since all predictions are identical, adjusting the threshold only has two possible outcomes:

  • If your threshold is less than or equal to 0.542: Every prediction will be TRUE (because 0.542 > threshold), so your confusion matrix will show all samples predicted as positive (like your current output where the TRUE column has all 58+46=104 cases, and the FALSE column is all 0).
  • If your threshold is greater than 0.542: Every prediction will be FALSE, so your confusion matrix will flip to show all samples predicted as negative (the FALSE column will have all 104 cases, and the TRUE column is 0).

If you're only adjusting thresholds within one of these ranges (e.g., testing 0.4, 0.5, 0.54), you won't see any change in the confusion matrix — because all predictions still fall on the same side of the threshold.

3. How to verify this

Run this quick check to confirm:

# Check the proportion of positive cases in your training set
mean(train$bigFire)

This should return exactly 0.542. Then try setting a threshold above that value, like 0.6:

CM <- table(test$bigFire, nullModel.pred > 0.6)
print(CM)

You'll see the confusion matrix completely change to all FALSE predictions.

4. Next steps if you want threshold-dependent performance

The null model is just a baseline — it doesn't use any of your features (month, FFMC, temp, etc.). To get predictions that vary across samples (and thus see confusion matrix changes with different thresholds), you'll need to fit a logistic regression model that includes predictors, e.g.:

fullModel <- glm(bigFire ~ month + FFMC + temp + RH, data = train, family = binomial)
fullModel.pred <- predict(fullModel, test, type = "response")

Now fullModel.pred will have different probabilities for different samples, so adjusting the threshold will change which samples are classified as positive/negative, and thus alter the confusion matrix.

内容的提问来源于stack exchange,提问作者Francesco De Santis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 15:32:47