You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Bootstrap重采样、LASSO与逐步回归的特征选择复现咨询

解决方案:手动控制Bootstrap重采样流程,实现特征选择

核心思路

由于boot库默认不会保存所有重采样的索引,导致无法复用同一组Bootstrap样本完成后续步骤。最直接的解决方式是手动生成并保存所有1000次重采样的索引,分阶段完成LASSO特征筛选和反向逐步Logistic回归,完全可控整个流程。


步骤1:生成并保存Bootstrap重采样索引

设置随机种子保证可复现,手动生成1000次重采样的行索引:

set.seed(123)
n_boot <- 1000
n_sample <- nrow(scaled_train)

# 生成1000次重采样的索引列表,每个元素对应一次Bootstrap样本的行号
boot_indices <- replicate(n_boot, sample(n_sample, replace = TRUE), simplify = FALSE)

步骤2:通过LASSO筛选Top10高频特征

遍历所有Bootstrap样本,运行LASSO并统计特征出现频率:

library(glmnet)

# 存储每次LASSO选中的非零特征
selected_features_lasso <- list()

for (i in 1:n_boot) {
  # 提取当前Bootstrap样本
  boot_sample <- scaled_train[boot_indices[[i]], ]
  y <- boot_sample$Outcome
  x_mat <- as.matrix(subset(boot_sample, select = -Outcome))
  
  # 交叉验证选择最优lambda
  cv_fit <- cv.glmnet(x_mat, y, family = "binomial", alpha = 1, standardize = TRUE)
  # 用最优lambda拟合LASSO
  lasso_fit <- glmnet(x_mat, y, family = "binomial", alpha = 1, lambda = cv_fit$lambda.min, standardize = TRUE)
  
  # 提取非零系数对应的特征(排除截距项)
  non_zero_mask <- coef(lasso_fit)[-1, ] != 0
  selected_features_lasso[[i]] <- names(non_zero_mask[non_zero_mask])
}

# 统计特征出现频率,选出Top10最常见的特征
feature_counts <- table(unlist(selected_features_lasso))
top10_features <- names(sort(feature_counts, decreasing = TRUE)[1:10])

步骤3:在同一Bootstrap样本上运行反向逐步Logistic回归

复用之前保存的索引,用Top10特征拟合反向逐步回归,统计最常见的特征组合:

# 存储每次逐步回归选中的特征
selected_features_stepwise <- list()

for (i in 1:n_boot) {
  # 提取同一Bootstrap样本,仅保留Top10特征和结局变量
  boot_sample <- scaled_train[boot_indices[[i]], ]
  data_subset <- boot_sample[, c(top10_features, "Outcome")]
  
  # 拟合全特征模型
  full_model <- glm(Outcome ~ ., data = data_subset, family = binomial)
  # 反向逐步回归(基于AIC筛选,关闭日志输出)
  step_model <- step(full_model, direction = "backward", trace = 0)
  
  # 提取最终模型的特征(排除截距项)
  selected_features_stepwise[[i]] <- attr(terms(step_model), "term.labels")
}

# 统计特征组合的出现频率,选出最常见的模型
feature_combinations <- sapply(selected_features_stepwise, function(x) paste(sort(x), collapse = ", "))
combination_counts <- table(feature_combinations)
most_common_features <- names(sort(combination_counts, decreasing = TRUE)[1])

备选方案:基于boot库保存索引

如果一定要使用boot库,可以修改自定义函数,让它同时返回LASSO系数和重采样索引:

lasso_Select <- function(x, indices){ 
  x_boot <- x[indices,]
  y <- x_boot$Outcome
  x_mat <- as.matrix(subset(x_boot, select = -Outcome))
  
  cv_fit <- cv.glmnet(x_mat, y, family = "binomial", alpha = 1, standardize = TRUE)
  lasso_fit <- glmnet(x_mat, y, family = "binomial", alpha = 1, lambda = cv_fit$lambda.min, standardize = TRUE)
  
  # 返回系数和当前重采样索引
  return(list(coef = coef(lasso_fit)[,1], indices = indices))
}

# 运行Bootstrap
myBootstrap <- boot(scaled_train, lasso_Select, R = 1000, parallel = "multicore", ncpus=5)

# 提取所有重采样索引
boot_indices <- lapply(myBootstrap$t, function(x) x$indices)
# 后续特征统计和逐步回归步骤与之前一致

不过手动生成索引的方式更直观,避免了boot库复杂的返回结构处理,推荐优先使用。

内容的提问来源于stack exchange,提问作者BenjaminHunter

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 09:27:54