You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言波士顿数据集递增训练集迭代线性回归指导请求

Hey there! Nice work getting the initial sampling code down—let’s break down the rest of the process step by step to help you run those 10 linear regression experiments and create your plot. Here's how to do it:

Step 1: Load Dependencies & Finalize Train/Test Splits

First, make sure you have the MASS package loaded (it contains the Boston dataset), then use your sampling index to split the data into fixed training and test sets. This way, your test set stays consistent across all experiments, which makes your results comparable:

# Load the MASS package for the Boston dataset
library(MASS)

# Your existing sampling code
sample <- sample.int(n = nrow(Boston), size = floor(.70*nrow(Boston)), replace = FALSE)

# Split into full training and fixed test sets
train_full <- Boston[sample, ]
test_set <- Boston[-sample, ]
Step 2: Prepare a Results Container

We’ll create a data frame to store each training set percentage and its corresponding test set RSE. This makes it easy to track and plot results later:

# Create empty data frame to hold experiment results
experiment_results <- data.frame(
  training_percent = seq(10, 100, 10),  # 10% to 100% in 10% increments
  test_rse = NA                         # Placeholder for RSE values
)
Step 3: Run the 10 Regression Experiments

Now loop through each training percentage, subset the full training set, train a linear regression model, predict on the test set, and calculate the test set RSE. I’m assuming your target variable is medv (median home value, the standard target for the Boston dataset)—adjust this if you’re using a different variable:

# Loop through each percentage in our results frame
for (i in 1:nrow(experiment_results)) {
  # Get the current training percentage
  current_pct <- experiment_results$training_percent[i]
  
  # Subset the full training set to the current percentage
  train_subset <- train_full[
    sample.int(n = nrow(train_full), 
               size = floor(current_pct / 100 * nrow(train_full)), 
               replace = FALSE),
    ]
  
  # Train linear regression model (predict medv using all other variables)
  model <- lm(medv ~ ., data = train_subset)
  
  # Generate predictions on the fixed test set
  test_predictions <- predict(model, newdata = test_set)
  
  # Calculate test set RSE (Root Squared Error)
  # This is the square root of the mean squared error between predictions and actual values
  test_rse <- sqrt(mean((test_set$medv - test_predictions)^2))
  
  # Store the calculated RSE in our results frame
  experiment_results$test_rse[i] <- test_rse
}

# Optional: Print results to verify everything worked
print(experiment_results)
Step 4: Plot Training Percentage vs Test RSE

You can use either base R or ggplot2 for plotting. Here are both options:

Option 1: Base R Plot (Quick & Simple)

plot(experiment_results$training_percent, experiment_results$test_rse,
     type = "b",  # Shows both line and data points
     col = "darkblue",
     lwd = 2,
     xlab = "Training Set Percentage",
     ylab = "Test Set RSE",
     main = "Test Set RSE vs Training Set Size")
grid()  # Add grid lines for better readability

Option 2: ggplot2 (Polished & Customizable)

If you don’t have ggplot2 installed, run install.packages("ggplot2") first:

library(ggplot2)

ggplot(experiment_results, aes(x = training_percent, y = test_rse)) +
  geom_line(color = "darkblue", linewidth = 1.2) +
  geom_point(size = 3, color = "darkred") +
  labs(
    x = "Training Set Percentage",
    y = "Test Set RSE",
    title = "How Training Set Size Impacts Test Set RSE"
  ) +
  theme_minimal() +
  theme(plot.title = element_text(hjust = 0.5))  # Center the plot title
Quick Notes to Keep in Mind
  • If your target variable isn’t medv, replace medv ~ . with your desired target (e.g., crim ~ . if predicting per capita crime rate).
  • Using a fixed test set eliminates variability from different train/test splits, making your comparison of training sizes more reliable.
  • If you meant the model's residual standard error instead of test set RMSE, let me know and I can adjust the calculation!

内容的提问来源于stack exchange,提问作者elle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:32:49