如何在R中利用研究起止日期按月聚合数据计算月度疾病患病率
队列研究月度患病率计算需求
我有队列研究数据,包含每位患者的研究起止日期,入组和退出时间各不相同。需要按月聚合数据,计算每个月的疾病患病率,输出需包含:
- 年月(
month_year) - 当月研究患者总数(
n_total):只要患者在该月有任何时间参与(包括当月退出的情况),都需计入 - 当月患病患者总数(
n_disease):仅统计当月处于患病状态的患者(即患病日期早于或等于当月,且患者当月在组) - 患病率(
prevalence):n_disease/n_total
注意:即使当月患者数为0、患病率为0,也必须保留对应行并标注患病率为0
当前数据格式
| patid | start_date | end_date | disease | disease_date |
|---|---|---|---|---|
| 1 | 01/03/2016 | 31/08/2021 | yes | 15/11/2017 |
| 2 | 24/03/2020 | 31/08/2021 | no | NA |
| 3 | 01/03/2020 | 23/08/2021 | yes | 15/08/2020 |
| 4 | 24/03/2016 | 01/08/2019 | no | NA |
| 5 | 24/03/2018 | 17/08/2020 | no | NA |
| 6 | 01/03/2016 | 04/08/2018 | yes | 01/01/2017 |
| 7 | 01/03/2016 | 31/08/2018 | yes | 18/03/2017 |
示例数据(R)
df <- data.frame(patid=c("1","2","3","4","5","6","7"), start_date=c("01/03/2016","24/08/2016", "01/01/2016","24/02/2016", "24/04/2016","01/04/2016", "01/09/2016"), end_date=c("31/12/2016","31/12/2016", "23/12/2016","01/08/2016", "17/06/2016","04/05/2016", "31/10/2016"), disease=c("yes","no","yes","no", "no","yes","yes"), disease_date=c("15/08/2016",NA, "15/08/2016",NA,NA, "01/05/2016","31/10/2016") )
解决方案(R实现)
步骤说明
- 转换日期格式为标准日期类型
- 生成研究时间范围内的所有月份序列
- 对每个患者,匹配其参与的所有月份
- 判断每个患者在对应月份是否处于患病状态
- 按月聚合统计总数和患病数,计算患病率
- 补全所有月份(包括无患者的月份)
代码实现
library(dplyr) library(lubridate) library(tidyr) # 1. 转换日期格式 df_clean <- df %>% mutate( start_date = dmy(start_date), end_date = dmy(end_date), disease_date = dmy(disease_date) ) # 2. 生成所有需要统计的月份序列 min_month <- floor_date(min(df_clean$start_date), "month") max_month <- floor_date(max(df_clean$end_date), "month") all_months <- tibble(month_year = seq(min_month, max_month, by = "month")) %>% mutate(month_year = format(month_year, "%m/%Y")) # 3. 扩展患者数据到每个参与的月份 patient_months <- df_clean %>% rowwise() %>% mutate( # 生成该患者参与的所有月份 months_in_study = list(seq(floor_date(start_date, "month"), floor_date(end_date, "month"), by = "month")) ) %>% unnest(months_in_study) %>% mutate(month_year = format(months_in_study, "%m/%Y")) %>% select(patid, month_year, disease, disease_date) # 4. 判断当月是否患病 patient_months <- patient_months %>% mutate( is_diseased = case_when( disease == "no" ~ FALSE, disease == "yes" & !is.na(disease_date) & disease_date <= months_in_study + months(1) - days(1) ~ TRUE, TRUE ~ FALSE ) ) # 5. 按月聚合统计 monthly_stats <- patient_months %>% group_by(month_year) %>% summarise( n_total = n(), n_disease = sum(is_diseased, na.rm = TRUE), .groups = "drop" ) %>% mutate(prevalence = ifelse(n_total == 0, 0, n_disease / n_total)) # 6. 补全所有月份,填充0值 final_result <- all_months %>% left_join(monthly_stats, by = "month_year") %>% replace_na(list(n_total = 0, n_disease = 0, prevalence = 0)) %>% mutate( n_total = as.character(n_total), n_disease = as.character(n_disease), prevalence = as.character(prevalence) ) # 查看结果 final_result
预期输出
structure(list(month_year = c("01/2016", "02/2016", "03/2016", "04/2016", "05/2016", "06/2016", "07/2016", "08/2016", "09/2016", "10/2016", "11/2016", "12/2016"), n_total = c("1", "2", "3", "5", "5", "4", "3", "4", "4", "4", "3", "3"), n_disease = c("0", "0", "0", "0", "1", "0", "0", "2", "0", "1", "0", "0"), prevalence = c("0", "0", "0", "0", "0.2", "0", "0", "0.5", "0", "0.25", "0", "0")), class = "data.frame", row.names = c(NA, -12L))
内容的提问来源于Stack Exchange,提问作者abrar_r
相关产品推荐
相关产品推荐

