You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言MASS包Boston数据集KNN(K=1-50)准确率绘图报错排查

Hey there! Let's break down what's going wrong with your code and fix those annoying NAs introduced by coercion warnings, plus get your KNN accuracy curve working properly.

Why the Warnings Happen

Your warning comes directly from this line:

sum(as.integer(ourBoston$high.medv))/length(result)

ourBoston$high.medv is a character vector with values "yes" and "no". When you try to convert these text values to integers with as.integer(), R has no way to map text to numbers, so it generates NAs for every entry—hence the repeated warning.

Other Critical Logic Issues

Beyond the warnings, your code has two key logic flaws:

  • You’re overwriting the result vector in every loop iteration: First you assign the KNN classification output (a character/factor vector) to result, then you try to set result[i] to a numeric accuracy value. This mixes data types and completely breaks your accuracy tracking.
  • Your "accuracy" calculation isn’t actually measuring model performance. Right now you’re just calculating the proportion of "yes" values in the original dataset, not how often the KNN model predicted correctly.

Fixed & Improved Code

Here’s the corrected version that resolves all these problems and produces the accuracy curve you want:

library(MASS)
library(class)  # Don't forget to load the `class` package for knn()!
library(ggplot2)

# Prepare the Boston dataset
ourBoston <- data.frame(Boston)
# Convert high.medv to a factor (optimal for classification tasks in R)
ourBoston$high.medv <- factor(ifelse(Boston$medv > 25, "yes", "no"), levels = c("no", "yes"))

# Initialize a numeric vector to store accuracy for each k
accuracy <- numeric(50)

# Loop through k values 1 to 50
for (i in 1:50) {
  # Run KNN model (using lstat as predictor, same data for train/test)
  knn_predictions <- knn(train = data.frame(ourBoston$lstat), 
                         test = data.frame(ourBoston$lstat), 
                         cl = ourBoston$high.medv, 
                         k = i)
  # Calculate actual accuracy: proportion of correct predictions
  accuracy[i] <- sum(knn_predictions == ourBoston$high.medv) / length(knn_predictions)
}

# Plot the accuracy curve
qplot(x = 1:50, y = accuracy, xlab = "Number of Neighbors (k)", ylab = "Accuracy", geom = "line") +
  ggtitle("KNN Accuracy vs. k (Boston Dataset)")

Key Changes Explained

  • Factor Conversion: Turning high.medv into a factor makes classification tasks like KNN work more reliably in R, and eliminates the need for problematic character-to-integer conversions.
  • Separate Prediction Storage: We use knn_predictions to store the model’s output, so we don’t overwrite our accuracy tracking vector.
  • Correct Accuracy Calculation: We directly compare predicted labels to true labels, then divide the number of matches by the total number of observations to get a valid accuracy score.
  • Explicit Package Loading: Added calls to load class (where knn() lives) and ggplot2 to avoid missing package errors.

Run this code, and you’ll get no more NA warnings—plus a proper plot showing how KNN accuracy changes as k increases from 1 to 50.

内容的提问来源于stack exchange,提问作者Akshay Lokhande

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:51:25