RandomForest调用报错:'User ID'未找到,如何忽略该列训练模型?
The error happens because your formula Purchased~. tells randomForest to use every column in the Network dataset as a predictor—including User ID. Since User ID has a space in its name, R can't parse it correctly without special handling, and anyway, we don't want this unique identifier cluttering up our model. Here are three straightforward fixes:
1. Modify the Formula to Exclude User ID
This is the simplest approach. Use the minus sign (-) to remove the User ID column from the predictor set, and wrap the column name in backticks because it contains a space:
library(randomForest) rfModel2 <- randomForest( formula = Purchased ~ . - `User ID`, data = Network, ntree = 50, importance = TRUE, replace = TRUE )
2. Subset the Dataset Before Training
If you prefer working with a stripped-down version of your data, you can subset the dataframe directly in the data argument to drop User ID:
library(randomForest) # Remove User ID by name rfModel2 <- randomForest( formula = Purchased~., data = Network[, !names(Network) %in% "User ID"], ntree = 50, importance = TRUE, replace = TRUE )
3. Use Tidyverse Syntax (If You're Comfortable With It)
If you use the dplyr package, you can use select() to cleanly drop the unwanted column:
library(randomForest) library(dplyr) rfModel2 <- randomForest( formula = Purchased~., data = Network %>% select(-`User ID`), ntree = 50, importance = TRUE, replace = TRUE )
All three methods will ensure your model only uses relevant predictors to train on the 1/0 Purchased target variable, while avoiding the missing object error from the spaced column name.
内容的提问来源于stack exchange,提问作者Denisse Kowaleski

