You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何创建适配regsubet、岭回归与套索回归的训练/测试集?

Creating Training/Test Sets for regsubsets, Ridge, and Lasso Regression

First off, your existing code for generating training/test indices works great as a foundation—it’s totally compatible across all three methods. Let’s start by restating that (with a little cleanup for readability):

# Load required package for the College dataset
library(ISLR)
set.seed(1) # Reproducibility is key!

# Generate training row indices (half the dataset, rounded up)
train_rows <- sample(nrow(College), ceiling(nrow(College)/2))

# Create logical indices for training and test sets
train_idx <- 1:nrow(College) %in% train_rows
test_idx <- !train_idx

The regsubsets Specific Tweaks

Here’s the thing about regsubsets (from the leaps package): it expects explicit separation of your feature matrix (X) and response variable (y), and it doesn’t play nicely with data frame subsetting in the same way lm or glmnet do. So we need to split our data into separate objects for features and response first.

Step 1: Split Features and Response

Let’s define our predictors and target variable for the College dataset (we’ll use Apps as the response here, but you can swap this for any column):

# Load the leaps package for regsubsets
library(leaps)

# Separate features (all columns except Apps) and response (Apps)
X <- College[, !names(College) %in% "Apps"]
y <- College$Apps

Step 2: Create Training/Test Datasets for regsubsets

Now we can use our logical indices to split both X and y into training and test sets. Note that regsubsets works with data frames for features, so we’ll keep X_train and X_test as data frames:

# Training sets
X_train <- X[train_idx, ]
y_train <- y[train_idx]

# Test sets
X_test <- X[test_idx, ]
y_test <- y[test_idx]

Step 3: Train a regsubsets Model

With these split datasets, you can train your subset selection model like this:

# Train regsubsets (e.g., exhaustive subset selection up to 10 features)
regsubsets_fit <- regsubsets(Apps ~ ., data = cbind(X_train, y_train), nvmax = 10)

Note: We use cbind(X_train, y_train) to recreate a data frame that includes the response variable, since regsubsets requires a formula interface with the data containing both predictors and response.

Compatibility with Ridge and Lasso

The same indices work perfectly for ridge/lasso (using glmnet). Since glmnet uses matrix inputs for features, we can convert our split data to matrices (though it also accepts data frames):

# Load glmnet for ridge/lasso
library(glmnet)

# Convert features to matrices (optional but standard for glmnet)
X_train_mat <- model.matrix(Apps ~ ., data = cbind(X_train, y_train))[, -1] # Remove intercept column
X_test_mat <- model.matrix(Apps ~ ., data = cbind(X_test, y_test))[, -1]

# Train ridge regression
ridge_fit <- glmnet(X_train_mat, y_train, alpha = 0)

# Train lasso regression
lasso_fit <- glmnet(X_train_mat, y_train, alpha = 1)

Key Takeaways

  • The core training/test index logic stays the same across all three methods—no need to reinvent the wheel there.
  • For regsubsets, explicitly split your features and response, then combine them back into a data frame when using the formula interface.
  • For glmnet, converting features to a matrix (with model.matrix to handle categorical variables) is a best practice, but the same indices work to split the data.

内容的提问来源于stack exchange,提问作者The Statistician Magician

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:57:13