R语言基于id、type和status条件过滤data.frame数据
解决方案
原代码问题说明
你之前的代码仅校验了type的序列连续性,完全没有涉及status字段的判断,也没有限制保留记录的范围到第一个type=2的位置,因此输出结果不符合预期。
实现逻辑
- 按id分组后,先排除没有type=2记录、或type=1记录不足2条的分组
- 定位每个分组第一次出现type=2的行位置,仅保留该位置及之前的所有记录
- 校验保留范围内所有type=1的记录的status是否全部为Unsuppressed,符合条件的分组才最终保留
完整代码
library(tidyverse) # 示例数据 data<- data.frame( id= c(215, 215, 215, 215, 297, 297, 297,297, 297,297,317,317,317,382,382,382,459,459,459), type=c(1,1,2,2,1,1,1,2,2,2,1,1,2,1,1,2,1,2,2), status=c("Unsuppressed","Unsuppressed","Unsuppressed","Unsuppressed","Unsuppressed","Suppressed","Unsuppressed","Unsuppressed","Suppressed", "Suppressed", "Unsuppressed", "Unsuppressed", "Unsuppressed", "Unsuppressed", "Unsuppressed", "Unsuppressed", "Unsuppressed", "Unsuppressed", "Suppressed") ) res <- data %>% group_by(id) %>% # 排除无type=2、或type=1不足2条的分组 filter(any(type == 2), sum(type == 1) >= 2) %>% # 定位分组内第一个type=2的位置 mutate(first_type2_idx = which(type == 2)[1]) %>% # 仅保留第一个type=2及之前的记录 filter(row_number() <= first_type2_idx) %>% # 校验所有type=1的status都是Unsuppressed filter(all(status[type == 1] == "Unsuppressed")) %>% ungroup() %>% # 移除辅助列 select(-first_type2_idx)
运行结果
> res # A tibble: 9 × 3 id type status <dbl> <dbl> <chr> 1 215 1 Unsuppressed 2 215 1 Unsuppressed 3 215 2 Unsuppressed 4 317 1 Unsuppressed 5 317 1 Unsuppressed 6 317 2 Unsuppressed 7 382 1 Unsuppressed 8 382 1 Unsuppressed 9 382 2 Unsuppressed
和你期望的输出完全一致。
内容的提问来源于stack exchange,提问作者Yebelay Berehan
相关产品推荐
相关产品推荐

