R语言:如何更简洁筛选含匹配项分组内的非匹配行?
R简洁实现:按ID组保留匹配行或全组行
需求说明
处理文件A(预期名称)与文件B(候选名称)的条目映射时,需实现以下逻辑:
- 若同一ID组内存在
status == "match"的行,仅保留该组内的匹配行,移除所有review行 - 若ID组内无匹配行,则保留该组所有行
示例数据
library(tidyverse) example <- tribble( ~id, ~expect, ~possible, ~status, 1, "box of forks", "spoon drawer", "review", 1, "box of forks", "box of forks", "match", 1, "box of forks", "cheese knife", "review", 2, "dish washer", "dish washing machine", "review", 2, "dish washer", "oven", "review", 2, "dish washer", "microwave", "review", )
当前实现(可行但繁琐)
example %>% group_by(id) %>% mutate(matches = case_when( status == "review" ~ 0, status == "match" ~ 1, ), total = sum(matches) ) %>% filter( !(matches == 0 & total > 0) )
简洁优化方案
方案1:无需额外列的极简写法
直接利用分组后的逻辑判断完成过滤,无需新增辅助列:
example %>% group_by(id) %>% filter(status == "match" | !any(status == "match")) %>% ungroup()
方案2:保留辅助列的简化写法
如果需要保留原结果中的matches和total列,可简化mutate逻辑:
example %>% group_by(id) %>% mutate( matches = as.integer(status == "match"), # 直接将逻辑值转为0/1,替代case_when total = sum(matches) ) %>% filter(status == "match" | total == 0) %>% # 更直观的过滤条件 ungroup()
预期输出
两种方案均可得到如下结果(方案2会保留matches和total列):
id expect possible status matches total <dbl> <chr> <chr> <chr> <dbl> <dbl> 1 1 box of forks box of forks match 1 1 2 2 dish washer dish washing machine review 0 0 3 2 dish washer oven review 0 0 4 2 dish washer microwave review 0 0
内容的提问来源于stack exchange,提问作者David Robie
相关产品推荐
相关产品推荐

