如何在mlr中为学习器封装指定特征选择?及多特征集逻辑回归重采样对比
Got it, let's break this down step by step using mlr in R—this is exactly the kind of workflow I’ve used repeatedly for feature subset comparisons. We’ll start with the core performance comparison, then wrap up how to encapsulate that feature selection logic directly into a learner.
Step 1: Setup Tasks & Define Feature Subsets
First, let’s assume you have your dataset my_data with a binary target column named "target". We’ll create the full task, then generate reduced tasks for each feature subset you want to test.
# Load mlr library(mlr) # Create full classification task full_task <- makeClassifTask(id = "full_features", data = my_data, target = "target") # Define your target feature subsets (example with 2 distinct groups) feat_subset1 <- c("age", "income", "gender") feat_subset2 <- c("age", "credit_score", "employment_status") # Generate reduced tasks (keep only your target features) reduced_task1 <- dropFeatures(full_task, setdiff(getTaskFeatureNames(full_task), feat_subset1)) reduced_task2 <- dropFeatures(full_task, setdiff(getTaskFeatureNames(full_task), feat_subset2))
Quick tip: setdiff(getTaskFeatureNames(full_task), feat_subsetX) grabs all features not in your target subset, so dropFeatures removes those—leaving exactly the features you want.
Step 2: Run Resampled Models & Compare Performance
Next, we’ll define our logistic regression learner, set up a robust resampling strategy, and run it on each task to get reliable performance estimates.
# Define logistic regression learner (binomial family for binary classification) lr_learner <- makeLearner("classif.logreg", predict.type = "prob") # Set resampling strategy (10-fold cross-validation, repeated 3 times for stability) resampling_strategy <- makeResampleDesc("RepCV", folds = 10, reps = 3) # Run resampling on each task res_full <- resample(lr_learner, full_task, resampling_strategy, measures = list(auc, acc)) res_sub1 <- resample(lr_learner, reduced_task1, resampling_strategy, measures = list(auc, acc)) res_sub2 <- resample(lr_learner, reduced_task2, resampling_strategy, measures = list(auc, acc)) # Print aggregated results for comparison cat("Full Features Performance:\n") print(res_full$aggr) cat("\nSubset 1 Performance:\n") print(res_sub1$aggr) cat("\nSubset 2 Performance:\n") print(res_sub2$aggr)
This gives you aggregated metrics (AUC and accuracy here) across all resampling iterations, so you can directly compare how each feature subset performs.
Step 3: Encapsulate Feature Selection Logic into a Learner
If you want to bundle the "use only X features" logic directly into a learner (so you don’t have to manually create reduced tasks every time), use a PreprocWrapper to add feature selection as a built-in preprocessing step.
Here’s how to create a custom learner that automatically uses your specified feature subset:
# Define a helper function to select your target features select_features <- function(data, target, args) { # Keep only the specified features plus the target column keep_cols <- c(args$feat_subset, target) data <- data[, keep_cols, drop = FALSE] return(list(data = data, control = list())) } # Create a preprocessing wrapper tied to your logistic regression learner feat_select_wrapper <- makePreprocWrapper( learner = lr_learner, train = select_features, predict = select_features, # Reuse the same logic for prediction par.vals = list(feat_subset = feat_subset1) # Pass your feature subset here ) # Now you can use this wrapper directly on the full task res_wrapper <- resample(feat_select_wrapper, full_task, resampling_strategy, measures = list(auc)) print(res_wrapper$aggr)
To create learners for other subsets, just make additional wrappers with different par.vals$feat_subset values—this way, feature selection is baked into the learner itself.
Bonus: Automate for Multiple Feature Subsets
If you have lots of subsets to test, loop through them to avoid repetitive code:
# List all feature subsets you want to evaluate all_subsets <- list( subset1 = c("age", "income", "gender"), subset2 = c("age", "credit_score", "employment_status"), subset3 = c("income", "credit_score", "gender") ) # Run resampling for each subset and store results results <- lapply(names(all_subsets), function(subset_name) { subset_feats <- all_subsets[[subset_name]] reduced_task <- dropFeatures(full_task, setdiff(getTaskFeatureNames(full_task), subset_feats)) res <- resample(lr_learner, reduced_task, resampling_strategy, measures = list(auc)) return(list(subset = subset_name, performance = res$aggr)) }) # Convert results to a clean data frame for easy comparison results_df <- do.call(rbind, lapply(results, function(x) { data.frame( Subset = x$subset, AUC = x$performance[[1]], stringsAsFactors = FALSE ) })) print(results_df)
内容的提问来源于stack exchange,提问作者Richi W

