如何在R中编写函数实现患者数据分组统计与宽表转换
解决方案
可以通过分组统计+宽格式转换实现需求,下面提供两种自动化实现方式,兼容多患者ID的场景:
方法1:使用tidyverse工具链(推荐)
借助dplyr和tidyr包实现简洁的数据流操作:
# 加载所需工具包 library(dplyr) library(tidyr) # 定义转换函数 transform_patient_data <- function(df) { df %>% # 按患者ID、单元类型、状态分组,统计每组记录数 group_by(Patient_ID, Unit_Type, Status) %>% summarise(count = n(), .groups = "drop") %>% # 将单元类型与状态合并为新列名 unite(col = "category", Unit_Type, Status, sep = "_") %>% # 转换为宽格式,缺失组合填充为0 pivot_wider(names_from = category, values_from = count, values_fill = 0) } # 测试函数 df <- structure(list(Patient_ID = c("1234", "1234", "1234", "1234", "1234", "1234", "1234", "1234", "1234"), Unit_Type = c("ABC", "ABC", "ABC", "ABC", "ABC", "DEF", "DEF", "DEF", "GHI"), Status = c("Returned", "Returned", "Returned", "Returned", "Transfused", "Transfused", "Transfused", "Transfused", "Transfused")), class = "data.frame", row.names = c(NA, -9L)) result <- transform_patient_data(df) print(result)
运行后输出结果:
| Patient_ID | ABC_Returned | ABC_Transfused | DEF_Transfused | GHI_Transfused |
|---|---|---|---|---|
| 1234 | 4 | 1 | 3 | 1 |
方法2:基础R实现(无依赖)
如果不想使用第三方包,可通过基础R函数完成转换:
transform_patient_data_base <- function(df) { # 生成三维交叉统计表 tab <- table(df[c("Patient_ID", "Unit_Type", "Status")]) # 转换为数据框格式 tab_df <- as.data.frame(tab, responseName = "count") # 合并单元类型与状态为列名前缀 tab_df$category <- paste(tab_df$Unit_Type, tab_df$Status, sep = "_") # 重塑为宽格式 reshape(tab_df, idvar = "Patient_ID", timevar = "category", direction = "wide", v.names = "count") } # 测试函数 result_base <- transform_patient_data_base(df) print(result_base)
函数特性
- 支持多患者ID批量处理,自动按患者单独统计
- 缺失的
单元类型-状态组合会自动填充0(tidyverse方法),避免NA值 - 列名自动按
Unit_Type_Status格式生成,完全匹配需求
内容的提问来源于stack exchange,提问作者Natasha H
相关产品推荐
相关产品推荐

