在R中添加数据集新行并预测信用评级:分类/回归选择及实操
Hey there! Let's work through your questions with clear examples and straightforward explanations.
First, let's fix a small oversight in your existing predict code—you forgot to reference your trained tree (train_tree) in the call! The correct syntax for predicting on your test set should be:
predict_test <- predict(train_tree, newdata = test, type = "class")
For a single customer, the process is nearly identical: you just need to package their data into a data frame that matches the exact structure (column names and data types) of your training set. Here's a concrete example:
Suppose your training data uses features like age, annual_income, credit_utilization, and payment_history (each with their respective scores). You'd create a single-row data frame for your new customer like this:
# Create a single-row data frame (match all feature names from your training data!) new_customer <- data.frame( age = 42, annual_income = 95000, credit_utilization = 0.3, payment_history = 85 ) # Use your pre-trained decision tree to predict their rating customer_rating <- predict(train_tree, newdata = new_customer, type = "class") # View the result cat("Predicted credit rating for the customer:", customer_rating, "\n")
Just make sure every feature from your training set is present in this new data frame—even if some values are missing (you’ll need to handle missing data the same way you did for training!).
The short, definitive answer: stick with classification for this task. Here’s why:
- Credit ratings are discrete categorical labels (e.g., "A", "B", "C", or "Good"/"Bad")—this is exactly what classification models are built to predict.
- Regression models are designed for continuous numerical values (like predicting loan amounts or monthly income), which doesn’t align with your goal of assigning a categorical rating.
When you add new rows to your dataset for prediction, just ensure:
- The new rows have the exact same feature names as your training data.
- Each feature uses the same data type (e.g., if
agewas numeric in training, it can’t be a character string in the new rows). - Your target variable (
rating) doesn’t need to be present in the new rows (since that’s what you’re trying to predict!).
One quick sanity check: Before training your decision tree, confirm that rating is treated as a factor (categorical variable) in your training data. If not, convert it with:
train$rating <- factor(train$rating)
This ensures rpart uses classification mode (which you already specified with method = "class", so you’re probably good to go here!).
内容的提问来源于stack exchange,提问作者AdeeThyag

