使用R包mRMRe从集成结果中筛选最优特征索引的技术咨询
Got it, let's work through how to pull out the highest-frequency feature indices from your mRMRe ensemble output. Your features_indices data frame has each ensemble run's selected features split across columns—we just need to tally up how often each index shows up across all runs.
Here are two straightforward approaches, depending on whether you prefer base R or the tidyverse:
Approach 1: Base R (No Extra Packages Needed)
# Step 1: Flatten the data frame into a single vector of all feature indices all_features <- unlist(features_indices) # Optional: Remove any NA values if your data has truncated/missing entries all_features <- na.omit(all_features) # Step 2: Count frequency of each index freq_table <- table(all_features) # Step 3: Convert to a sorted data frame for readability freq_df <- as.data.frame(freq_table) colnames(freq_df) <- c("feature_index", "frequency") # Rename columns for clarity freq_df$feature_index <- as.numeric(as.character(freq_df$feature_index)) # Convert factor to numeric # Step 4: Sort by frequency (highest first) sorted_freq <- freq_df[order(-freq_df$frequency), ] # Example: Get the top 10 most frequent features top_10_features <- sorted_freq$feature_index[1:10] # Print the sorted results print(sorted_freq)
Approach 2: Tidyverse (dplyr + tidyr)
If you're comfortable with the tidyverse, this is a more concise workflow:
library(tidyverse) # Reshape data to long format, count frequencies, and sort sorted_freq <- features_indices %>% pivot_longer(everything(), names_to = "ensemble_run", values_to = "feature_index") %>% drop_na(feature_index) %>% # Remove missing values count(feature_index, sort = TRUE) %>% # Count and sort by frequency rename(frequency = n) # Rename count column # Example: Get features that appear in at least 3 ensemble runs high_freq_features <- sorted_freq %>% filter(frequency >= 3) %>% pull(feature_index) # Print the results print(sorted_freq)
Quick Note on Your Sample Data
Looking at the snippet you shared, indices like 1406 and 2709 show up in all 5 runs (frequency = 5), so they'll land at the top of your sorted list. You can decide how many top features to keep—either pick a fixed number (like top 10) or set a minimum frequency threshold (e.g., features that appear in ≥3 runs).
内容的提问来源于stack exchange,提问作者Shuvayan Das

