如何按规则筛选DataFrame的Top 10样本?含分组限制与NA处理
Let's break down how to achieve your desired result step by step. First, let's recap the rules to make sure we're aligned:
- Keep all samples with
group = NAthat fall into the top value range - For samples with non-NA groups, only keep 1 sample per group (the one with the highest value)
- First retain the specific samples: 36, 22, 2, 20, 11, then fill the remaining 5 spots with the next highest-value samples that follow the rules.
Step 1: Reproduce the Original Data
First, let's generate the dataset as you provided:
set.seed(1) df <- data.frame( sample = 1:50, value = runif(50), group = c(rep(NA, 20), gl(3, 10)) )
Step 2: Define Priority Samples & Sort the Data
First, list the samples you want to prioritize, then sort the entire dataset by value in descending order (so we're always picking from the highest values first):
library(dplyr) # Define the samples we need to keep first priority_samples <- c(36, 22, 2, 20, 11) # Sort the full dataset by value descending df_sorted <- df %>% arrange(desc(value))
Step 3: Extract Priority Samples
Pull out the priority samples first—these are guaranteed to be in our final result:
df_priority <- df_sorted %>% filter(sample %in% priority_samples)
Step 4: Process Remaining Samples (Follow Rules)
For the remaining samples, we need to apply the group constraint:
- Keep all NA-group samples (since they don't have group limits)
- For non-NA groups, only keep the highest-value sample per group (since we already sorted,
slice(1)gives us the top value for each group)
df_remaining <- df_sorted %>% filter(!sample %in% priority_samples) %>% # Exclude already kept samples group_by(group) %>% slice(1) %>% # Keep top 1 per group (NA groups are kept as-is) ungroup()
Step 5: Fill the Remaining Spots
We need 5 more samples to reach 10 total. Grab the top 5 from our processed remaining samples:
df_additional <- df_remaining %>% slice_head(n = 5)
Step 6: Combine & Finalize
Merge the priority samples and additional samples, then re-sort by value to get the final ordered result:
df_final <- bind_rows(df_priority, df_additional) %>% arrange(desc(value)) %>% slice_head(n = 10) # Ensure we have exactly 10 samples # View the final result print(df_final)
Alternative One-Liner Approach
If you prefer a more condensed version, you can mark priority samples and combine the logic:
df_final <- df %>% arrange(desc(value)) %>% mutate(is_priority = sample %in% priority_samples) %>% # First take all priority samples filter(is_priority) %>% bind_rows( # Then take top 1 per group from non-priority samples df %>% arrange(desc(value)) %>% filter(!sample %in% priority_samples) %>% group_by(group) %>% slice(1) %>% ungroup() ) %>% arrange(desc(value)) %>% slice_head(n = 10)
This will give you exactly the 10 samples you want: the 5 priority ones, plus the next 5 highest-value samples that adhere to the group uniqueness rule.
内容的提问来源于stack exchange,提问作者user42485

