葡萄酒数据集Logistic回归测试集预测结果缺失排查、混淆矩阵构建及阈值转换代码影响咨询
1. Why your prediction results are shorter than the test set?
Looking at your code, the immediate culprit is a variable name typo:
You assigned the prediction output to result1, but then used an undefined variable result in the ifelse line:
result1 = predict(wmodel, newdata = test, type = 'response') # Oops! Should use result1 here, not result result1 = ifelse(result > 0.5, 1, 0)
If there was an existing variable named result in your environment with 393 entries, this line would overwrite result1 with that shorter dataset's converted values—explaining why you only get 393 predictions instead of 480.
To confirm and fix this:
- First, check the length of your raw prediction output before the
ifelsestep: runlength(result1)right afterpredict()—it should matchnrow(test)(480) - Correct the
ifelseline to useresult1instead ofresult - Double-check that your
testdataset has all the same variables astrain(glm will drop rows from newdata if there are missing variables, but you mentioned no NAs, so this is less likely here)
2. What happens if you remove the ifelse line?
Removing that line means result1 will stay as the raw predicted probabilities from the logistic regression—values between 0 and 1 representing the model's confidence that each sample is in the "Good" class (since family=binomial predicts the probability of the positive class).
Here's what that changes:
- You can't directly build a confusion matrix with
test$final_takeyet, because confusion matrices require hard 0/1 labels, not probabilities - You gain flexibility: you can experiment with different classification thresholds (not just 0.5) to adjust precision/recall, or calculate metrics like AUC-ROC that use the full probability distribution
- From your existing stats, it looks like your model is predicting mostly "Bad" (0) at the 0.5 threshold—keeping the raw probabilities lets you see exactly how confident the model is about each prediction
Fixed code snippet to get matching predictions and confusion matrix
n=nrow(wine_log) shuffled=wine_log[sample(n),] train_indices=1:round(0.7*n) test_indices=(round(0.7*n)+1):n train=shuffled[train_indices,] test=shuffled[test_indices,] wmodel = glm(final_take~., family = binomial, data=train) summary(wmodel) # Get predictions and fix the variable name typo result1 = predict(wmodel, newdata = test, type = 'response') result1 = ifelse(result1 > 0.5, 1, 0) # Now uses the correct variable # Verify lengths match cat("Test set rows:", nrow(test), "\nPrediction length:", length(result1), "\n") # Build confusion matrix confusion_matrix = table(Predicted = result1, Actual = test$final_take) print(confusion_matrix)
内容的提问来源于stack exchange,提问作者Bhavna

