R语言中按因子水平出现次数筛选data.frame列的方法
Hey folks, let's solve this specific R problem where you need to clean up a data frame with two-level factor columns—keeping only those columns where both factor levels show up more than once, and dropping any columns where one level appears ≤1 times (we'll also retain non-factor columns like row IDs by default).
Example Data
First, let's set up the sample data frame we're working with:
df <- data.frame( row.name = 1:4, Factor1 = c("dog", "dog", "dog", "dog"), Factor2 = c("dog", "dog", "cat", "cat"), Factor3 = c("cat", "cat", "dog", "dog"), Factor4 = c("cat", "dog", "dog", "dog"), stringsAsFactors = TRUE )
In this example, we want to retain Factor2 and Factor3 (both levels appear twice), while dropping Factor1 (only "dog" exists) and Factor4 ("cat" only shows up once).
Working Solutions
After tweaking a method from Ryan, here are two solid approaches that get the job done perfectly:
1. Using dplyr (Tidyverse Style)
If you prefer the tidyverse syntax, use select_if to filter columns conditionally:
library(dplyr) filtered_df <- df %>% select_if(~ !is.factor(.x) | sum(tabulate(.x) > 1) >= 2)
How it works:
- For non-factor columns (like
row.name), we keep them automatically with!is.factor(.x). - For factor columns,
tabulate(.x)counts how many times each level appears. We then sum how many levels have a count greater than 1—if that sum is ≥2, both levels meet our requirement, so we keep the column.
2. Using Base R
If you want to stick to base R without extra packages, use sapply to create a boolean subset vector:
filtered_df <- df[, sapply(df, function(x) !is.factor(x) | sum(tabulate(.x) > 1) >= 2)]
How it works:
- This uses the exact same logic as the dplyr version.
sapplyiterates over each column, runs the check, and returnsTRUE/FALSEfor each column. We use that vector to keep only the columns where the check passes.
Result
Either method will give you the filtered data frame you need:
row.name Factor2 Factor3 1 1 dog cat 2 2 dog cat 3 3 cat dog 4 4 cat dog
内容的提问来源于stack exchange,提问作者plik

