波士顿房价数据集线性回归报错:MEDV未找到但数据集含该字段
Hey there, I get it—you’re scratching your head because you know the MEDV column exists in your dataset, but R keeps throwing that frustrating error when you run your lm() call. Let’s break down what’s going on and fix this step by step:
1. Double-check your column names (hidden characters are sneaky!)
The error mentions CAT..MEDV, which suggests your target column might have been renamed accidentally (maybe during data import or a previous operation). First, confirm exactly what columns are in your dataset with these quick checks:
# Print all column names to spot discrepancies colnames(data) # Or glance at the first few rows to verify head(data)
If you see CAT..MEDV instead of MEDV, just update your model formula to match the actual column name:
mlr1 <- lm(CAT..MEDV ~ CRIM + CHAS + RM, data = data.training)
2. Verify your training set includes the target column
It’s rare, but sometimes sampling can go wonky (though your code looks correct at first glance). Let’s confirm MEDV made it into your training split:
# Check if MEDV is in the training set columns "MEDV" %in% colnames(data.training) # Inspect the training set directly head(data.training)
If the result is FALSE, re-run your sampling code—maybe a random quirk caused the column to be excluded (unlikely, but worth ruling out).
3. Confirm your dataset loaded correctly
If you imported the Boston Housing data from a CSV or external file, double-check your import code. For example, if you used read.csv() without header=TRUE, R might have treated your column names as the first data row. Fix that with:
# Re-import with header enabled (adjust the file path as needed) data <- read.csv("boston_housing.csv", header = TRUE)
4. Refresh the BostonHousing dataset (if using the mlbench version)
If you’re using the built-in BostonHousing dataset from the mlbench package, it’s possible a previous operation modified the data. Reset it with:
# Install and load mlbench if you haven't already install.packages("mlbench") library(mlbench) # Reload the fresh dataset data(BostonHousing) data <- BostonHousing # Re-run your full workflow n.training <- floor(nrow(data)*0.7) id.training <- sample(1:nrow(data), n.training) data.training <- data[id.training,] data.test <- data[-id.training,] mlr1 <- lm(MEDV ~ CRIM + CHAS + RM, data = data.training)
One of these steps should get your model running smoothly!
内容的提问来源于stack exchange,提问作者pepapalacios

