R语言:按ID分组计算特定条件下daystoevent差值并新增列
R语言分组计算指定事件的天数差值
针对每个ID分组,我们需要新增一列,计算该组中第一个excision == 1对应的daystoevent与died == 1对应的daystoevent的差值;若分组中不存在died == 1或excision == 1的记录,差值设为NA。
方法一:使用dplyr包(推荐,语法简洁直观)
先加载dplyr包,通过分组后计算目标值并生成新列:
library(dplyr) df_result <- df %>% group_by(ID) %>% mutate( # 获取分组内第一个excision=1的daystoevent first_excision_day = first(daystoevent[excision == 1]), # 获取分组内died=1的daystoevent(假设每组最多1条死亡记录) died_day = daystoevent[died == 1], # 计算差值 day_diff = died_day - first_excision_day ) %>% ungroup() # 输出结果 print(df_result)
运行结果:
# A tibble: 10 × 6 ID daystoevent excision died first_excision_day died_day day_diff <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> 1 1 19116 0 0 19201 NA NA 2 1 19201 1 0 19201 NA NA 3 1 19399 0 0 19201 NA NA 4 1 19416 0 0 19201 NA NA 5 2 17017 1 0 17017 17085 68 6 2 17036 0 0 17017 17085 68 7 2 17085 0 1 17017 17085 68 8 3 17017 1 0 17017 17085 68 9 3 17036 1 0 17017 17085 68 10 3 17085 0 1 17017 17085 68
方法二:使用Base R(无需额外包)
通过拆分数据框、分组计算再合并的方式实现:
# 按ID拆分数据框 df_split <- split(df, df$ID) # 对每个分组进行计算 df_list <- lapply(df_split, function(x) { # 取第一个excision=1的daystoevent first_excision <- x$daystoevent[x$excision == 1][1] # 取died=1的daystoevent died_day <- x$daystoevent[x$died == 1] # 计算差值,无对应记录则设为NA x$day_diff <- if (length(died_day) > 0 && length(first_excision) > 0) { died_day - first_excision } else { NA } return(x) }) # 合并回数据框并重置行名 df_result_base <- do.call(rbind, df_list) rownames(df_result_base) <- NULL # 输出结果 print(df_result_base)
两种方法最终结果一致,可根据自身习惯选择使用。
内容的提问来源于stack exchange,提问作者ltong
相关产品推荐
相关产品推荐

