You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用survSplit()按事件时间拆分复发事件生存数据集?

复发事件数据下用survSplit()按个体拆分生存数据集的问题

问题背景

在单个体对应单条记录的生存数据场景中,可轻松将数据转换为计数过程形式——按所有受试者的事件时间拆分个体观测时长(不超过个体实际观测周期),单个个体对应多条记录。但在复发事件数据场景下(单个个体有多条记录对应多次事件,最后一条记录可能为事件或截尾),直接使用survival包的survSplit()函数时,发现它是按行(而非按ID)扩展数据,无法实现仅在个体内按全局事件时间拆分的需求。

示例代码

library(survival)
library(dplyr)

dat <-  structure(list(ID = c(1L, 1L, 1L, 1L, 2L, 2L, 2L, 3L, 3L, 3L, 3L), 
                       age = c(43L, 43L, 43L, 43L, 43L, 43L, 43L, 41L, 41L, 41L, 41L), 
                       treat = structure(c(2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), 
                                         levels = c("old", "new"), class = "factor"), 
                       time0 = c(0L, 6L, 9L, 56L, 0L, 42L, 87L, 0L, 15L, 17L, 36L), 
                       time1 = c(6L, 9L, 56L, 88L, 42L, 87L, 91L, 15L, 17L, 36L, 112L), 
                       status = c(1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 0), 
                       event = c(1L, 2L, 3L, 4L, 1L, 2L, 3L, 1L, 2L, 3L, 4L)),
                  datalabel = "Chapter 9 Exercises", time.stamp = " 7 Dec 1999 08:53", formats = c("%9.0g", "%9.0g", "%9.0g", "%9.0g", "%9.0g", "%19.0g", "%9.0g"), 
                  types = c(105L, 98L, 98L, 105L, 105L, 98L, 98L), 
                  val.labels = c("", "", "oldnew", "", "", "censor", ""), 
                  var.labels = c("Subject Identification", "Age", "Treatment Assignment", "Time of Last Episode", "Time of Current Episode or censoring", "Indicator for Soreness Episode or censoring", "Soreness Episode Number"), 
                  row.names = c("3", "4", "1", "2", "5", "7", "6", "8", "9", "10", "11"), 
                  version = 6L, label.table = list(oldnew = structure(0:1, names = c("new", "old")), 
                                                   censor = structure(0:1, names = c("censored", "experienced"))), class = "data.frame")

# 提取全局事件时间
event_times <- sort(unique(with(dat, time1[status == 1])))
# 直接使用survSplit得到的结果不符合需求:按行扩展而非按ID拆分
dat2 <- survSplit(Surv(time1, status) ~., dat, cut = event_times)

解决方案

survSplit()本身不支持直接对复发事件数据按ID拆分,因为它的设计逻辑是处理每行对应的生存区间,而非追踪单个个体的完整时间线。需要先将个体的多条记录整理为连续的时间区间序列,再使用survSplit()处理:

步骤1:整理个体的连续时间区间

先将每个个体的所有事件/截尾时间整合为连续的时间区间,并匹配对应状态与协变量:

# 提取每个个体的所有时间点并排序
id_time_points <- dat %>%
  group_by(ID) %>%
  summarise(times = sort(unique(c(time0, time1))), .groups = "drop")

# 生成每个个体的连续时间区间,并匹配对应状态
id_intervals <- id_time_points %>%
  rowwise() %>%
  mutate(
    start = head(times, -1),
    end = tail(times, -1),
    status = map2(start, end, ~dat$status[dat$ID == ID & dat$time0 == .x])
  ) %>%
  unnest(c(start, end, status)) %>%
  select(ID, start, end, status)

# 合并个体的协变量信息
id_full <- id_intervals %>%
  left_join(dat %>% distinct(ID, age, treat), by = "ID")

步骤2:用survSplit()按全局事件时间拆分

基于整理后的个体连续区间数据集,使用survSplit()实现按个体内时间拆分:

# 按全局事件时间拆分,生成计数过程形式的数据集
dat_cp <- survSplit(Surv(start, end, status) ~ ., data = id_full, cut = event_times)

这样得到的dat_cp就是按个体拆分的计数过程数据集,每个个体的观测时间被全局事件时间拆分,每条记录对应一个时间段及该段内的事件状态。

内容的提问来源于stack exchange,提问作者LucaS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 10:25:54