R语言波士顿数据集递增训练集迭代线性回归指导请求
Hey there! Nice work getting the initial sampling code down—let’s break down the rest of the process step by step to help you run those 10 linear regression experiments and create your plot. Here's how to do it:
First, make sure you have the MASS package loaded (it contains the Boston dataset), then use your sampling index to split the data into fixed training and test sets. This way, your test set stays consistent across all experiments, which makes your results comparable:
# Load the MASS package for the Boston dataset library(MASS) # Your existing sampling code sample <- sample.int(n = nrow(Boston), size = floor(.70*nrow(Boston)), replace = FALSE) # Split into full training and fixed test sets train_full <- Boston[sample, ] test_set <- Boston[-sample, ]
We’ll create a data frame to store each training set percentage and its corresponding test set RSE. This makes it easy to track and plot results later:
# Create empty data frame to hold experiment results experiment_results <- data.frame( training_percent = seq(10, 100, 10), # 10% to 100% in 10% increments test_rse = NA # Placeholder for RSE values )
Now loop through each training percentage, subset the full training set, train a linear regression model, predict on the test set, and calculate the test set RSE. I’m assuming your target variable is medv (median home value, the standard target for the Boston dataset)—adjust this if you’re using a different variable:
# Loop through each percentage in our results frame for (i in 1:nrow(experiment_results)) { # Get the current training percentage current_pct <- experiment_results$training_percent[i] # Subset the full training set to the current percentage train_subset <- train_full[ sample.int(n = nrow(train_full), size = floor(current_pct / 100 * nrow(train_full)), replace = FALSE), ] # Train linear regression model (predict medv using all other variables) model <- lm(medv ~ ., data = train_subset) # Generate predictions on the fixed test set test_predictions <- predict(model, newdata = test_set) # Calculate test set RSE (Root Squared Error) # This is the square root of the mean squared error between predictions and actual values test_rse <- sqrt(mean((test_set$medv - test_predictions)^2)) # Store the calculated RSE in our results frame experiment_results$test_rse[i] <- test_rse } # Optional: Print results to verify everything worked print(experiment_results)
You can use either base R or ggplot2 for plotting. Here are both options:
Option 1: Base R Plot (Quick & Simple)
plot(experiment_results$training_percent, experiment_results$test_rse, type = "b", # Shows both line and data points col = "darkblue", lwd = 2, xlab = "Training Set Percentage", ylab = "Test Set RSE", main = "Test Set RSE vs Training Set Size") grid() # Add grid lines for better readability
Option 2: ggplot2 (Polished & Customizable)
If you don’t have ggplot2 installed, run install.packages("ggplot2") first:
library(ggplot2) ggplot(experiment_results, aes(x = training_percent, y = test_rse)) + geom_line(color = "darkblue", linewidth = 1.2) + geom_point(size = 3, color = "darkred") + labs( x = "Training Set Percentage", y = "Test Set RSE", title = "How Training Set Size Impacts Test Set RSE" ) + theme_minimal() + theme(plot.title = element_text(hjust = 0.5)) # Center the plot title
- If your target variable isn’t
medv, replacemedv ~ .with your desired target (e.g.,crim ~ .if predicting per capita crime rate). - Using a fixed test set eliminates variability from different train/test splits, making your comparison of training sizes more reliable.
- If you meant the model's residual standard error instead of test set RMSE, let me know and I can adjust the calculation!
内容的提问来源于stack exchange,提问作者elle

