如何用survSplit()按事件时间拆分复发事件生存数据集?
复发事件数据下用survSplit()按个体拆分生存数据集的问题
问题背景
在单个体对应单条记录的生存数据场景中,可轻松将数据转换为计数过程形式——按所有受试者的事件时间拆分个体观测时长(不超过个体实际观测周期),单个个体对应多条记录。但在复发事件数据场景下(单个个体有多条记录对应多次事件,最后一条记录可能为事件或截尾),直接使用survival包的survSplit()函数时,发现它是按行(而非按ID)扩展数据,无法实现仅在个体内按全局事件时间拆分的需求。
示例代码
library(survival) library(dplyr) dat <- structure(list(ID = c(1L, 1L, 1L, 1L, 2L, 2L, 2L, 3L, 3L, 3L, 3L), age = c(43L, 43L, 43L, 43L, 43L, 43L, 43L, 41L, 41L, 41L, 41L), treat = structure(c(2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), levels = c("old", "new"), class = "factor"), time0 = c(0L, 6L, 9L, 56L, 0L, 42L, 87L, 0L, 15L, 17L, 36L), time1 = c(6L, 9L, 56L, 88L, 42L, 87L, 91L, 15L, 17L, 36L, 112L), status = c(1, 1, 1, 1, 1, 1, 0, 1, 1, 1, 0), event = c(1L, 2L, 3L, 4L, 1L, 2L, 3L, 1L, 2L, 3L, 4L)), datalabel = "Chapter 9 Exercises", time.stamp = " 7 Dec 1999 08:53", formats = c("%9.0g", "%9.0g", "%9.0g", "%9.0g", "%9.0g", "%19.0g", "%9.0g"), types = c(105L, 98L, 98L, 105L, 105L, 98L, 98L), val.labels = c("", "", "oldnew", "", "", "censor", ""), var.labels = c("Subject Identification", "Age", "Treatment Assignment", "Time of Last Episode", "Time of Current Episode or censoring", "Indicator for Soreness Episode or censoring", "Soreness Episode Number"), row.names = c("3", "4", "1", "2", "5", "7", "6", "8", "9", "10", "11"), version = 6L, label.table = list(oldnew = structure(0:1, names = c("new", "old")), censor = structure(0:1, names = c("censored", "experienced"))), class = "data.frame") # 提取全局事件时间 event_times <- sort(unique(with(dat, time1[status == 1]))) # 直接使用survSplit得到的结果不符合需求:按行扩展而非按ID拆分 dat2 <- survSplit(Surv(time1, status) ~., dat, cut = event_times)
解决方案
survSplit()本身不支持直接对复发事件数据按ID拆分,因为它的设计逻辑是处理每行对应的生存区间,而非追踪单个个体的完整时间线。需要先将个体的多条记录整理为连续的时间区间序列,再使用survSplit()处理:
步骤1:整理个体的连续时间区间
先将每个个体的所有事件/截尾时间整合为连续的时间区间,并匹配对应状态与协变量:
# 提取每个个体的所有时间点并排序 id_time_points <- dat %>% group_by(ID) %>% summarise(times = sort(unique(c(time0, time1))), .groups = "drop") # 生成每个个体的连续时间区间,并匹配对应状态 id_intervals <- id_time_points %>% rowwise() %>% mutate( start = head(times, -1), end = tail(times, -1), status = map2(start, end, ~dat$status[dat$ID == ID & dat$time0 == .x]) ) %>% unnest(c(start, end, status)) %>% select(ID, start, end, status) # 合并个体的协变量信息 id_full <- id_intervals %>% left_join(dat %>% distinct(ID, age, treat), by = "ID")
步骤2:用survSplit()按全局事件时间拆分
基于整理后的个体连续区间数据集,使用survSplit()实现按个体内时间拆分:
# 按全局事件时间拆分,生成计数过程形式的数据集 dat_cp <- survSplit(Surv(start, end, status) ~ ., data = id_full, cut = event_times)
这样得到的dat_cp就是按个体拆分的计数过程数据集,每个个体的观测时间被全局事件时间拆分,每条记录对应一个时间段及该段内的事件状态。
内容的提问来源于stack exchange,提问作者LucaS
相关产品推荐
相关产品推荐

