You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何筛选分组内同时满足双条件的id对应数据行

R按分组条件筛选数据行实现方法

需求说明

现有名为df的数据框对象,相同id属于同一分组,需要筛选出同时满足以下两个条件的id对应的所有数据行:

  • 同一分组下的code字段取值包含以大写字母I开头的内容,不限制I后续跟随的数字,例如I11、I31这类取值均满足条件
  • 同一分组下的code字段包含精确取值E12

示例数据中符合筛选要求的id为1和2,二者分组内同时存在含大写I开头的编码以及E12编码,示例数据结构与打印结果如下:

structure(list(id = c(1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3, 
3, 3, 3, 4, 4, 4, 4, 4, 4), diag = c("main", "other", "main", 
"other", "main", "other", "main", "other", "main", "other", "main", 
"other", "main", "other", "main", "other", "main", "other", "main", 
"other", "main", "other"), code = c("I11", "E12", "I11", "Q34", 
"I31", "C33", "E12", "I34", "E12", "I45", "E12", "Z11", "E13", 
"Z12", "E14", "Z13", "I25", "E1", "I25", "E2", "I25", "E3")), class = c("grouped_df", 
"tbl_df", "tbl", "data.frame"), row.names = c(NA, -22L), groups = structure(list(
    id = c(1, 2, 3, 4), .rows = structure(list(1:6, 7:10, 11:16, 
        17:22), ptype = integer(0), class = c("vctrs_list_of", 
    "vctrs_vctr", "list"))), class = c("tbl_df", "tbl", "data.frame"
), row.names = c(NA, -4L), .drop = TRUE))

> df
# A tibble: 22 × 3
# Groups:   id [4]
      id diag  code 
   <dbl> <chr> <chr>
 1     1 main  I11  
 2     1 other E12  
 3     1 main  I11  
 4     1 other Q34  
 5     1 main  I31  
 6     1 other C33  
 7     2 main  E12  
 8     2 other I34  
 9     2 main  E12  
10     2 other I45  
# … with 12 more rows

实现代码

直接使用dplyr包的分组筛选逻辑即可,由于原数据本身已经是按id分组的tibble对象,不需要重复执行分组操作:

library(dplyr)

res <- df %>%
  filter(
    # 判断组内是否存在以I开头的code
    any(grepl("^I", code)),
    # 判断组内是否存在精确等于E12的code
    any(code == "E12")
  )

逻辑说明

  • grepl("^I", code)通过正则表达式匹配字符串开头为大写I的取值,自动覆盖I后接任意数字的场景,不会误匹配字符串中间带I的编码
  • any()函数会在当前id分组范围内判断是否至少有1行满足对应条件,两个判断条件同时成立时,会保留该分组下的所有行
  • 示例数据运行后返回id为1、2的共10行数据,和预期结果一致

内容的提问来源于stack exchange,提问作者zhiwei li

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 11:24:39