R语言MASS包Boston数据集KNN(K=1-50)准确率绘图报错排查
Hey there! Let's break down what's going wrong with your code and fix those annoying NAs introduced by coercion warnings, plus get your KNN accuracy curve working properly.
Why the Warnings Happen
Your warning comes directly from this line:
sum(as.integer(ourBoston$high.medv))/length(result)
ourBoston$high.medv is a character vector with values "yes" and "no". When you try to convert these text values to integers with as.integer(), R has no way to map text to numbers, so it generates NAs for every entry—hence the repeated warning.
Other Critical Logic Issues
Beyond the warnings, your code has two key logic flaws:
- You’re overwriting the
resultvector in every loop iteration: First you assign the KNN classification output (a character/factor vector) toresult, then you try to setresult[i]to a numeric accuracy value. This mixes data types and completely breaks your accuracy tracking. - Your "accuracy" calculation isn’t actually measuring model performance. Right now you’re just calculating the proportion of
"yes"values in the original dataset, not how often the KNN model predicted correctly.
Fixed & Improved Code
Here’s the corrected version that resolves all these problems and produces the accuracy curve you want:
library(MASS) library(class) # Don't forget to load the `class` package for knn()! library(ggplot2) # Prepare the Boston dataset ourBoston <- data.frame(Boston) # Convert high.medv to a factor (optimal for classification tasks in R) ourBoston$high.medv <- factor(ifelse(Boston$medv > 25, "yes", "no"), levels = c("no", "yes")) # Initialize a numeric vector to store accuracy for each k accuracy <- numeric(50) # Loop through k values 1 to 50 for (i in 1:50) { # Run KNN model (using lstat as predictor, same data for train/test) knn_predictions <- knn(train = data.frame(ourBoston$lstat), test = data.frame(ourBoston$lstat), cl = ourBoston$high.medv, k = i) # Calculate actual accuracy: proportion of correct predictions accuracy[i] <- sum(knn_predictions == ourBoston$high.medv) / length(knn_predictions) } # Plot the accuracy curve qplot(x = 1:50, y = accuracy, xlab = "Number of Neighbors (k)", ylab = "Accuracy", geom = "line") + ggtitle("KNN Accuracy vs. k (Boston Dataset)")
Key Changes Explained
- Factor Conversion: Turning
high.medvinto a factor makes classification tasks like KNN work more reliably in R, and eliminates the need for problematic character-to-integer conversions. - Separate Prediction Storage: We use
knn_predictionsto store the model’s output, so we don’t overwrite our accuracy tracking vector. - Correct Accuracy Calculation: We directly compare predicted labels to true labels, then divide the number of matches by the total number of observations to get a valid accuracy score.
- Explicit Package Loading: Added calls to load
class(whereknn()lives) andggplot2to avoid missing package errors.
Run this code, and you’ll get no more NA warnings—plus a proper plot showing how KNN accuracy changes as k increases from 1 to 50.
内容的提问来源于stack exchange,提问作者Akshay Lokhande

