使用dplyr::filter多列迭代筛选含至少一个符合条件元素的行
Got it, let's work through this problem! You want to use dplyr::filter() to keep only rows where at least one column meets your target condition, and you need to iterate across multiple columns to do it. Here's how to make this work with your dataset dat:
Step 1: Set Up Your Data
First, let's formalize the dataset you provided (I filled in the truncated Date.8 column to keep consistency):
library(dplyr) # Define your dataset dat <- structure(list( Date.1 = c(NA, NA, NA, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7), Date.2 = c(NA, NA, NA, 7, 7, 6, 6, 6, 6, 6, 6, 6, 6, 6, 6), Date.3 = c(NA, NA, NA, 6, 6, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8), Date.4 = c(NA, NA, NA, 8, 8, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7), Date.5 = c(NA, NA, NA, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7 ), Date.6 = c(NA, NA, NA, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7), Date.7 = c(NA, NA, NA, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7 ), Date.8 = c(NA, NA, NA, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7, 7) ))
Step 2: Use if_any() for Column Iteration
The cleanest and most efficient way (available in dplyr 1.0.0+) is to use if_any(), which is built exactly for this kind of "check any column" task.
Example: Filter rows where at least one column equals 7
Let's say your condition is "value equals 7" (adjust this to your actual condition):
# Keep rows where at least one Date.* column is 7 filtered_dat <- dat %>% filter(if_any(starts_with("Date."), ~ .x == 7))
Breakdown of the code:
starts_with("Date."): Targets all columns that start with "Date." (useeverything()to check all columns, orc(Date.1, Date.3)to specify exact columns)~ .x == 7: The condition to check—replace==7with your actual logic (e.g.,!is.na(.x)to keep rows with at least one non-NA value,.x > 6to check for values greater than 6)
Alternative for Older dplyr Versions
If you're using a dplyr version before 1.0.0, use rowwise() + c_across() to check each row:
# Compatible with older dplyr versions filtered_dat_old <- dat %>% rowwise() %>% filter(any(c_across(starts_with("Date.")) == 7, na.rm = TRUE)) %>% ungroup()
na.rm = TRUE: Ensures NA values don't break theany()check (remove this if you want NA to count as a non-matching value)
Result
Running either of these will filter out the first 3 rows (all NA values, no matches) and keep the remaining rows, since every row from 4 onwards has at least one column with a value of 7.
内容的提问来源于stack exchange,提问作者zirodec

