You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言生存分析数据重构正确性验证问询

R语言医疗患者数据重构与时变生存分析适配问题

原始数据与预处理

我正在用R处理医疗患者数据,原始数据构建及预处理代码如下:

my_data = data.frame(id = c(1,2,3), status_2017 = c("alive", "alive", "alive"), status_2018 = c("alive", "dead", "alive"), status_2019 = c("alive", "dead", "dead"), height_2017 = rnorm(3,3,3), height_2018 = rnorm(3,3,3), 
                     height_2019 = rnorm(3,3,3) , weight_2017  = rnorm(3,3,3), weight_2018 = rnorm(3,3,3), weight_2019 = rnorm(3,3,3))

cols <- colnames(my_data)
ix <- my_data[, startsWith(cols, "status")] == "dead"

my_data[, startsWith(cols, "height")][ ix ] <- NA
my_data[, startsWith(cols, "weight")][ ix ] <- NA

预处理后的数据示例:

id status_2017 status_2018 status_2019 height_2017 height_2018 height_2019 weight_2017 weight_2018 weight_2019
1  1       alive       alive       alive   3.7276706    4.524869   -1.648458   -1.702781    7.755581    3.369895
2  2       alive        dead        dead   0.7539518          NA          NA    1.060408          NA          NA
3  3       alive       alive        dead   6.6213771    2.122374          NA    5.114120    1.851467          NA

数据重构需求

我需要将数据重构为以下格式:

  • 每位患者每年对应一行
  • 新增“year”列
  • 将status_2017、status_2018、status_2019合并为单个“status”列
  • 将height_2017、height_2018、height_2019合并为单个“height”列
  • 将weight_2017、weight_2018、weight_2019合并为单个“weight”列
  • 创建新变量new_var:若患者ID存在2019年数据,则所有行new_var为0;其他患者的非最大年度行new_var为0,最大年度行new_var为1

我的实现代码与结果

我尝试用以下代码实现:

library(dplyr)
library(tidyr)

my_data_long <- na.omit(my_data %>%
    pivot_longer(cols = -c(id, status_2017),
                 names_to = c(".value", "year"),
                 names_pattern = "(height|weight)_(\d{4})") %>%
    arrange(id, year))


final = my_data_long  %>%
  group_by(id) %>%
  mutate(
    new_var = ifelse(any(year == "2019"), 0, 1),
    max_year = max(year)
  ) %>%
  ungroup() %>%
  mutate(
    new_var = ifelse(year == max_year & new_var == 1, 1, 0),
    max_year = NULL
  )

得到的结果:

> final
# A tibble: 6 x 6
     id status_2017 year  height weight new_var
  <dbl> <chr>       <chr>  <dbl>  <dbl>   <dbl>
1     1 alive       2017   2.39    2.27       0
2     1 alive       2018  -0.541   1.63       0
3     1 alive       2019  -1.93   10.1        0
4     2 alive       2017   4.18   -3.35       1
5     3 alive       2017  -1.35    7.12       0
6     3 alive       2018   1.42    1.70       1

疑问与目标

我的最终目标是重构数据以适配时变生存分析模型(如Cox-PH),请问当前的实现是否正确?

时间差添加尝试

另外,我尝试为每个ID添加时间差,代码及结果如下:

library(stringr)

final %>%
  group_by(id) %>%
  mutate(start = 0:(n() - 1),
         end = 1:n()) %>%
  ungroup()

# A tibble: 6 x 8
     id status_2017 year  height weight new_var start   end
  <dbl> <chr>       <chr>  <dbl>  <dbl>   <dbl> <int> <int>
1     1 alive       2017   2.39    2.27       0     0     1
2     1 alive       2018  -0.541   1.63       0     1     2
3     1 alive       2019  -1.93   10.1        0     2     3
4     2 alive       2017   4.18   -3.35       1     0     1
5     3 alive       2017  -1.35    7.12       0     0     1
6     3 alive       2018   1.42    1.70       1     1     2

内容的提问来源于stack exchange,提问作者stats_noob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 06:24:35