如何创建适配regsubet、岭回归与套索回归的训练/测试集?
regsubsets, Ridge, and Lasso Regression First off, your existing code for generating training/test indices works great as a foundation—it’s totally compatible across all three methods. Let’s start by restating that (with a little cleanup for readability):
# Load required package for the College dataset library(ISLR) set.seed(1) # Reproducibility is key! # Generate training row indices (half the dataset, rounded up) train_rows <- sample(nrow(College), ceiling(nrow(College)/2)) # Create logical indices for training and test sets train_idx <- 1:nrow(College) %in% train_rows test_idx <- !train_idx
The regsubsets Specific Tweaks
Here’s the thing about regsubsets (from the leaps package): it expects explicit separation of your feature matrix (X) and response variable (y), and it doesn’t play nicely with data frame subsetting in the same way lm or glmnet do. So we need to split our data into separate objects for features and response first.
Step 1: Split Features and Response
Let’s define our predictors and target variable for the College dataset (we’ll use Apps as the response here, but you can swap this for any column):
# Load the leaps package for regsubsets library(leaps) # Separate features (all columns except Apps) and response (Apps) X <- College[, !names(College) %in% "Apps"] y <- College$Apps
Step 2: Create Training/Test Datasets for regsubsets
Now we can use our logical indices to split both X and y into training and test sets. Note that regsubsets works with data frames for features, so we’ll keep X_train and X_test as data frames:
# Training sets X_train <- X[train_idx, ] y_train <- y[train_idx] # Test sets X_test <- X[test_idx, ] y_test <- y[test_idx]
Step 3: Train a regsubsets Model
With these split datasets, you can train your subset selection model like this:
# Train regsubsets (e.g., exhaustive subset selection up to 10 features) regsubsets_fit <- regsubsets(Apps ~ ., data = cbind(X_train, y_train), nvmax = 10)
Note: We use
cbind(X_train, y_train)to recreate a data frame that includes the response variable, sinceregsubsetsrequires a formula interface with the data containing both predictors and response.
Compatibility with Ridge and Lasso
The same indices work perfectly for ridge/lasso (using glmnet). Since glmnet uses matrix inputs for features, we can convert our split data to matrices (though it also accepts data frames):
# Load glmnet for ridge/lasso library(glmnet) # Convert features to matrices (optional but standard for glmnet) X_train_mat <- model.matrix(Apps ~ ., data = cbind(X_train, y_train))[, -1] # Remove intercept column X_test_mat <- model.matrix(Apps ~ ., data = cbind(X_test, y_test))[, -1] # Train ridge regression ridge_fit <- glmnet(X_train_mat, y_train, alpha = 0) # Train lasso regression lasso_fit <- glmnet(X_train_mat, y_train, alpha = 1)
Key Takeaways
- The core training/test index logic stays the same across all three methods—no need to reinvent the wheel there.
- For
regsubsets, explicitly split your features and response, then combine them back into a data frame when using the formula interface. - For
glmnet, converting features to a matrix (withmodel.matrixto handle categorical variables) is a best practice, but the same indices work to split the data.
内容的提问来源于stack exchange,提问作者The Statistician Magician

