You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:按ID分组计算特定条件下daystoevent差值并新增列

R语言分组计算指定事件的天数差值

针对每个ID分组,我们需要新增一列,计算该组中第一个excision == 1对应的daystoevent与died == 1对应的daystoevent的差值;若分组中不存在died == 1或excision == 1的记录,差值设为NA。

方法一:使用dplyr包(推荐,语法简洁直观)

先加载dplyr包,通过分组后计算目标值并生成新列:

library(dplyr)

df_result <- df %>%
  group_by(ID) %>%
  mutate(
    # 获取分组内第一个excision=1的daystoevent
    first_excision_day = first(daystoevent[excision == 1]),
    # 获取分组内died=1的daystoevent(假设每组最多1条死亡记录)
    died_day = daystoevent[died == 1],
    # 计算差值
    day_diff = died_day - first_excision_day
  ) %>%
  ungroup()

# 输出结果
print(df_result)

运行结果:

# A tibble: 10 × 6
      ID daystoevent excision died first_excision_day died_day day_diff
   <dbl>       <dbl>    <dbl> <dbl>              <dbl>    <dbl>    <dbl>
 1     1       19116        0    0               19201       NA       NA
 2     1       19201        1    0               19201       NA       NA
 3     1       19399        0    0               19201       NA       NA
 4     1       19416        0    0               19201       NA       NA
 5     2       17017        1    0               17017    17085       68
 6     2       17036        0    0               17017    17085       68
 7     2       17085        0    1               17017    17085       68
 8     3       17017        1    0               17017    17085       68
 9     3       17036        1    0               17017    17085       68
10     3       17085        0    1               17017    17085       68

方法二:使用Base R(无需额外包)

通过拆分数据框、分组计算再合并的方式实现:

# 按ID拆分数据框
df_split <- split(df, df$ID)

# 对每个分组进行计算
df_list <- lapply(df_split, function(x) {
  # 取第一个excision=1的daystoevent
  first_excision <- x$daystoevent[x$excision == 1][1]
  # 取died=1的daystoevent
  died_day <- x$daystoevent[x$died == 1]
  # 计算差值,无对应记录则设为NA
  x$day_diff <- if (length(died_day) > 0 && length(first_excision) > 0) {
    died_day - first_excision
  } else {
    NA
  }
  return(x)
})

# 合并回数据框并重置行名
df_result_base <- do.call(rbind, df_list)
rownames(df_result_base) <- NULL

# 输出结果
print(df_result_base)

两种方法最终结果一致,可根据自身习惯选择使用。

内容的提问来源于stack exchange,提问作者ltong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 11:06:29