如何合并dplyr::filter_at结果,实现多变量批量筛选全同值行?
Great question! Let's break down what's causing your unexpected results and walk through a few clean, scalable solutions that work even for dozens of columns.
First, let's clarify why your initial attempt with all_vars(. == 5 | . == 1) didn't work: that code checks if each individual column is either 1 or 5, but it doesn't enforce that all columns in the row are the same value. That's why you ended up with rows where values mix 1 and 5 (like row 2: A=1, B=5, C=5) — each column passes the "1 or 5" check, but they aren't all identical.
Solution 1: Use dplyr's modern across (recommended, since filter_at is soft-deprecated)
If you're using dplyr 1.0.0+, across is the preferred way to handle operations across multiple columns. Here's how to filter rows where all values in columns A-C are identical (regardless of what that value is):
library(dplyr) # First, recreate your data frame (fixed the D column to match your expected output) A <- c(1, 1, 1, 5, 5) B <- c(2, 5, 1, 5, 5) C <- c(3, 5, 1, 5, 5) D <- c("T", "F", "T", "T", "T") df <- data.frame(A, B, C, D) # Filter rows where all values in A-C are the same df %>% rowwise() %>% filter(all(c_across(A:C) == first(c_across(A:C)))) %>% ungroup()
This will give you exactly your desired output:
# A tibble: 3 × 4 A B C D <dbl> <dbl> <dbl> <chr> 1 1 1 1 T 2 5 5 5 T 3 5 5 5 T
If you only want to keep rows where the identical value is 1 or 5 (in case there are other identical values you don't care about), just add an extra filter:
df %>% rowwise() %>% filter(all(c_across(A:C) == first(c_across(A:C)))) %>% ungroup() %>% filter(A %in% c(1, 5)) # Since all columns are identical, checking A is enough
Solution 2: A base R-inspired approach (no rowwise needed)
If you prefer avoiding rowwise (which can be slower for large datasets), you can use rowSums to check if every column in the row matches the first column:
df %>% filter(rowSums(across(A:C) == pull(select(., A))) == ncol(select(., A:C)))
Solution 3: Chain multiple filter_at calls (your original idea)
If you specifically want to stick with filter_at (though it's not the latest dplyr syntax), you can combine multiple conditions with | to select rows that are all 1 or all 5:
df %>% filter( filter_at(vars(A:C), all_vars(. == 1)) | filter_at(vars(A:C), all_vars(. == 5)) )
This works, but it's less scalable if you need to check for more values later (you'd have to add more filter_at clauses). The first solution is better for flexibility.
Key Takeaway
The core issue with your initial code was mixing up "each column is X or Y" vs "all columns are the same value". The rowwise + c_across approach is clean, scalable, and works for any number of columns without manually listing each variable.
内容的提问来源于stack exchange,提问作者numeralpotatochips

