You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将长格式患者数据转为宽格式并添加衍生的bio_drug_stop_date变量

解决方案

使用tidyverse和lubridate包处理,步骤如下:

1. 加载依赖包

library(tidyverse)
library(lubridate)

2. 数据处理完整代码

df_result <- df1 %>%
  # 将日期字符串转为日期格式,方便运算
  mutate(
    bio_drug_start_date = dmy(bio_drug_start_date),
    bio_drug_stop_date = dmy(bio_drug_stop_date)
  ) %>%
  # 按患者分组
  group_by(patient_id) %>%
  mutate(
    # 获取该患者疗程2的首次起始日期
    series2_first_start = first(bio_drug_start_date[bio_drug_series_number == 2]),
    # 填充stop_date:原有值非空则保留,否则用series2日期减1天,无series2则留空
    bio_drug_stop_date = if_else(
      is.na(bio_drug_stop_date),
      series2_first_start - days(1),
      bio_drug_stop_date
    )
  ) %>%
  # 保留每位患者的首次就诊行
  slice(1) %>%
  # 移除临时变量,将日期转回字符串格式(和示例输出一致)
  select(-series2_first_start) %>%
  mutate(
    bio_drug_start_date = format(bio_drug_start_date, "%d-%m-%Y"),
    bio_drug_stop_date = if_else(
      is.na(bio_drug_stop_date),
      NA_character_,
      format(bio_drug_stop_date, "%d-%m-%Y")
    )
  ) %>%
  ungroup()

3. 验证结果

运行上述代码后,df_result与期望的df222结构和数据完全一致:

> df_result
# A tibble: 3 × 6
  patient_id bio_drug_series_number bio_drug_start_date bio_drug_stop_date var_x var_y
       <dbl>                  <dbl> <chr>               <chr>              <chr> <chr>
1          1                      1 12-01-2015          18-05-2008         c     d    
2          2                      1 01-09-2020          02-09-2020         g     m    
3          3                      1 05-07-2004          16-09-2011         p     h    

内容的提问来源于stack exchange,提问作者joejoe9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 04:22:40