R语言生存分析数据重构正确性验证问询
R语言医疗患者数据重构与时变生存分析适配问题
原始数据与预处理
我正在用R处理医疗患者数据,原始数据构建及预处理代码如下:
my_data = data.frame(id = c(1,2,3), status_2017 = c("alive", "alive", "alive"), status_2018 = c("alive", "dead", "alive"), status_2019 = c("alive", "dead", "dead"), height_2017 = rnorm(3,3,3), height_2018 = rnorm(3,3,3), height_2019 = rnorm(3,3,3) , weight_2017 = rnorm(3,3,3), weight_2018 = rnorm(3,3,3), weight_2019 = rnorm(3,3,3)) cols <- colnames(my_data) ix <- my_data[, startsWith(cols, "status")] == "dead" my_data[, startsWith(cols, "height")][ ix ] <- NA my_data[, startsWith(cols, "weight")][ ix ] <- NA
预处理后的数据示例:
id status_2017 status_2018 status_2019 height_2017 height_2018 height_2019 weight_2017 weight_2018 weight_2019 1 1 alive alive alive 3.7276706 4.524869 -1.648458 -1.702781 7.755581 3.369895 2 2 alive dead dead 0.7539518 NA NA 1.060408 NA NA 3 3 alive alive dead 6.6213771 2.122374 NA 5.114120 1.851467 NA
数据重构需求
我需要将数据重构为以下格式:
- 每位患者每年对应一行
- 新增“year”列
- 将status_2017、status_2018、status_2019合并为单个“status”列
- 将height_2017、height_2018、height_2019合并为单个“height”列
- 将weight_2017、weight_2018、weight_2019合并为单个“weight”列
- 创建新变量new_var:若患者ID存在2019年数据,则所有行new_var为0;其他患者的非最大年度行new_var为0,最大年度行new_var为1
我的实现代码与结果
我尝试用以下代码实现:
library(dplyr) library(tidyr) my_data_long <- na.omit(my_data %>% pivot_longer(cols = -c(id, status_2017), names_to = c(".value", "year"), names_pattern = "(height|weight)_(\d{4})") %>% arrange(id, year)) final = my_data_long %>% group_by(id) %>% mutate( new_var = ifelse(any(year == "2019"), 0, 1), max_year = max(year) ) %>% ungroup() %>% mutate( new_var = ifelse(year == max_year & new_var == 1, 1, 0), max_year = NULL )
得到的结果:
> final # A tibble: 6 x 6 id status_2017 year height weight new_var <dbl> <chr> <chr> <dbl> <dbl> <dbl> 1 1 alive 2017 2.39 2.27 0 2 1 alive 2018 -0.541 1.63 0 3 1 alive 2019 -1.93 10.1 0 4 2 alive 2017 4.18 -3.35 1 5 3 alive 2017 -1.35 7.12 0 6 3 alive 2018 1.42 1.70 1
疑问与目标
我的最终目标是重构数据以适配时变生存分析模型(如Cox-PH),请问当前的实现是否正确?
时间差添加尝试
另外,我尝试为每个ID添加时间差,代码及结果如下:
library(stringr) final %>% group_by(id) %>% mutate(start = 0:(n() - 1), end = 1:n()) %>% ungroup() # A tibble: 6 x 8 id status_2017 year height weight new_var start end <dbl> <chr> <chr> <dbl> <dbl> <dbl> <int> <int> 1 1 alive 2017 2.39 2.27 0 0 1 2 1 alive 2018 -0.541 1.63 0 1 2 3 1 alive 2019 -1.93 10.1 0 2 3 4 2 alive 2017 4.18 -3.35 1 0 1 5 3 alive 2017 -1.35 7.12 0 0 1 6 3 alive 2018 1.42 1.70 1 1 2
内容的提问来源于stack exchange,提问作者stats_noob
相关产品推荐
相关产品推荐

