R语言如何筛选分组内同时满足双条件的id对应数据行
R按分组条件筛选数据行实现方法
需求说明
现有名为df的数据框对象,相同id属于同一分组,需要筛选出同时满足以下两个条件的id对应的所有数据行:
- 同一分组下的
code字段取值包含以大写字母I开头的内容,不限制I后续跟随的数字,例如I11、I31这类取值均满足条件 - 同一分组下的
code字段包含精确取值E12
示例数据中符合筛选要求的id为1和2,二者分组内同时存在含大写I开头的编码以及E12编码,示例数据结构与打印结果如下:
structure(list(id = c(1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 4), diag = c("main", "other", "main", "other", "main", "other", "main", "other", "main", "other", "main", "other", "main", "other", "main", "other", "main", "other", "main", "other", "main", "other"), code = c("I11", "E12", "I11", "Q34", "I31", "C33", "E12", "I34", "E12", "I45", "E12", "Z11", "E13", "Z12", "E14", "Z13", "I25", "E1", "I25", "E2", "I25", "E3")), class = c("grouped_df", "tbl_df", "tbl", "data.frame"), row.names = c(NA, -22L), groups = structure(list( id = c(1, 2, 3, 4), .rows = structure(list(1:6, 7:10, 11:16, 17:22), ptype = integer(0), class = c("vctrs_list_of", "vctrs_vctr", "list"))), class = c("tbl_df", "tbl", "data.frame" ), row.names = c(NA, -4L), .drop = TRUE)) > df # A tibble: 22 × 3 # Groups: id [4] id diag code <dbl> <chr> <chr> 1 1 main I11 2 1 other E12 3 1 main I11 4 1 other Q34 5 1 main I31 6 1 other C33 7 2 main E12 8 2 other I34 9 2 main E12 10 2 other I45 # … with 12 more rows
实现代码
直接使用dplyr包的分组筛选逻辑即可,由于原数据本身已经是按id分组的tibble对象,不需要重复执行分组操作:
library(dplyr) res <- df %>% filter( # 判断组内是否存在以I开头的code any(grepl("^I", code)), # 判断组内是否存在精确等于E12的code any(code == "E12") )
逻辑说明
grepl("^I", code)通过正则表达式匹配字符串开头为大写I的取值,自动覆盖I后接任意数字的场景,不会误匹配字符串中间带I的编码any()函数会在当前id分组范围内判断是否至少有1行满足对应条件,两个判断条件同时成立时,会保留该分组下的所有行- 示例数据运行后返回id为1、2的共10行数据,和预期结果一致
内容的提问来源于stack exchange,提问作者zhiwei li
相关产品推荐
相关产品推荐

