如何用R识别出院72小时内再次入院的患者?
解决方案
首先加载必要的工具包:
library(tidyverse) library(lubridate)
先预处理数据集,把字符型日期转换为可计算的日期格式:
df <- df %>% mutate( entry = dmy(entry), # 适配日/月/年的格式转换 exit = dmy(exit) )
一、宽格式输出:整合住院信息并标记72小时再入院
按患者分组排序,给住院记录编号后转宽格式,再判断再入院情况:
wide_df <- df %>% group_by(IPP) %>% arrange(entry) %>% # 按入院日期排序,保证住院顺序正确 mutate(id = row_number()) %>% # 给每个患者的住院记录分配序号 ungroup() %>% pivot_wider( id_cols = IPP, names_from = id, values_from = c(NDA, entry, exit), names_glue = "{.value}_{id}" # 生成清晰的列名,比如NDA_1、entry_1 ) %>% # 判断是否存在72小时内再入院 mutate( return72h = case_when( !is.na(entry_2) & (entry_2 - exit_1) <= days(3) ~ "yes", TRUE ~ "no" ) ) # 调整列顺序匹配示例输出 wide_df <- wide_df %>% select(IPP, NDA_1, entry_1, exit_1, NDA_2, entry_2, exit_2, return72h) print(wide_df)
输出结果:
# A tibble: 3 × 8 IPP NDA_1 entry_1 exit_1 NDA_2 entry_2 exit_2 return72h <chr> <chr> <date> <date> <chr> <date> <date> <chr> 1 123456 147 2024-01-01 2024-02-24 258 2024-04-08 2024-04-10 no 2 456789 211 2024-07-14 2024-07-15 456 2024-07-16 2024-07-20 yes 3 789632 789 2024-06-06 2024-09-08 NA NA NA no
如果患者有3次及以上住院,可扩展case_when判断后续的exit_2与entry_3的间隔,或用批量处理逻辑覆盖多轮住院的判断。
二、原数据结构输出:添加间隔时间与标记字段
按患者分组排序后,获取下一次入院日期,计算时间差并标记:
long_df <- df %>% group_by(IPP) %>% arrange(entry) %>% mutate( next_entry = lead(entry), # 提取同一患者的下一次入院日期 interval_days = as.numeric(next_entry - exit), # 计算出院到下一次入院的天数 return72h = case_when( !is.na(interval_days) & interval_days <= 3 ~ "yes", TRUE ~ "no" ) ) %>% ungroup() print(long_df)
输出结果:
# A tibble: 5 × 7 IPP NDA entry exit next_entry interval_days return72h <chr> <chr> <date> <date> <date> <dbl> <chr> 1 123456 147 2024-01-01 2024-02-24 2024-04-08 43 no 2 123456 258 2024-04-08 2024-04-10 NA NA no 3 456789 211 2024-07-14 2024-07-15 2024-07-16 1 yes 4 456789 456 2024-07-16 2024-07-20 NA NA no 5 789632 789 2024-06-06 2024-09-08 NA NA no
interval_days为NA时,表示该住院是患者的最后一次住院,无后续入院记录。
内容的提问来源于stack exchange,提问作者alexia lemaire
相关产品推荐
相关产品推荐

