如何用R提取各id首行disease值为1的分组数据?
解决方法
给定数据集:
df<-data.frame(id=c(1,1,1,2,2,2,2,3,3), date=c(20220311,20220315,20220317,20220514,20220517,20220518,20220519,20220613,20220618), disease=c(0,1,0,1,1,1,0,1,1))
需求:提取所有首行disease值为1的id对应的全部数据。
方法1:使用dplyr(tidyverse生态)
library(dplyr) # 先筛选出符合条件的id target_ids <- df %>% group_by(id) %>% slice_head(n = 1) %>% # 获取每组第一行 filter(disease == 1) %>% pull(id) # 提取这些id的所有数据 result_df <- df %>% filter(id %in% target_ids)
方法2:Base R(无需额外安装包)
# 给每行标记对应id的首行disease值 first_disease_per_id <- ave(df$disease, df$id, FUN = function(x) x[1]) # 筛选首行disease为1的所有行 result_df <- df[first_disease_per_id == 1, ]
验证结果,result_df输出与预期一致:
> result_df id date disease 4 2 20220514 1 5 2 20220517 1 6 2 20220518 1 7 2 20220519 0 8 3 20220613 1 9 3 20220618 1
内容的提问来源于stack exchange,提问作者Lee
相关产品推荐
相关产品推荐

