You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R包mRMRe从集成结果中筛选最优特征索引的技术咨询

Extracting Most Frequent Feature Indices from mRMRe Ensemble Results

Got it, let's work through how to pull out the highest-frequency feature indices from your mRMRe ensemble output. Your features_indices data frame has each ensemble run's selected features split across columns—we just need to tally up how often each index shows up across all runs.

Here are two straightforward approaches, depending on whether you prefer base R or the tidyverse:

Approach 1: Base R (No Extra Packages Needed)

# Step 1: Flatten the data frame into a single vector of all feature indices
all_features <- unlist(features_indices)

# Optional: Remove any NA values if your data has truncated/missing entries
all_features <- na.omit(all_features)

# Step 2: Count frequency of each index
freq_table <- table(all_features)

# Step 3: Convert to a sorted data frame for readability
freq_df <- as.data.frame(freq_table)
colnames(freq_df) <- c("feature_index", "frequency")  # Rename columns for clarity
freq_df$feature_index <- as.numeric(as.character(freq_df$feature_index))  # Convert factor to numeric

# Step 4: Sort by frequency (highest first)
sorted_freq <- freq_df[order(-freq_df$frequency), ]

# Example: Get the top 10 most frequent features
top_10_features <- sorted_freq$feature_index[1:10]

# Print the sorted results
print(sorted_freq)

Approach 2: Tidyverse (dplyr + tidyr)

If you're comfortable with the tidyverse, this is a more concise workflow:

library(tidyverse)

# Reshape data to long format, count frequencies, and sort
sorted_freq <- features_indices %>%
  pivot_longer(everything(), names_to = "ensemble_run", values_to = "feature_index") %>%
  drop_na(feature_index) %>%  # Remove missing values
  count(feature_index, sort = TRUE) %>%  # Count and sort by frequency
  rename(frequency = n)  # Rename count column

# Example: Get features that appear in at least 3 ensemble runs
high_freq_features <- sorted_freq %>%
  filter(frequency >= 3) %>%
  pull(feature_index)

# Print the results
print(sorted_freq)

Quick Note on Your Sample Data

Looking at the snippet you shared, indices like 1406 and 2709 show up in all 5 runs (frequency = 5), so they'll land at the top of your sorted list. You can decide how many top features to keep—either pick a fixed number (like top 10) or set a minimum frequency threshold (e.g., features that appear in ≥3 runs).

内容的提问来源于stack exchange,提问作者Shuvayan Das

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:33:02