You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中编写函数实现患者数据分组统计与宽表转换

解决方案

可以通过分组统计+宽格式转换实现需求,下面提供两种自动化实现方式,兼容多患者ID的场景:

方法1:使用tidyverse工具链(推荐)

借助dplyr和tidyr包实现简洁的数据流操作:

# 加载所需工具包
library(dplyr)
library(tidyr)

# 定义转换函数
transform_patient_data <- function(df) {
  df %>%
    # 按患者ID、单元类型、状态分组,统计每组记录数
    group_by(Patient_ID, Unit_Type, Status) %>%
    summarise(count = n(), .groups = "drop") %>%
    # 将单元类型与状态合并为新列名
    unite(col = "category", Unit_Type, Status, sep = "_") %>%
    # 转换为宽格式,缺失组合填充为0
    pivot_wider(names_from = category, values_from = count, values_fill = 0)
}

# 测试函数
df <- structure(list(Patient_ID = c("1234", "1234", "1234", "1234", 
"1234", "1234", "1234", "1234", "1234"), Unit_Type = c("ABC", 
"ABC", "ABC", "ABC", "ABC", "DEF", "DEF", "DEF", "GHI"), Status = c("Returned", 
"Returned", "Returned", "Returned", "Transfused", "Transfused", 
"Transfused", "Transfused", "Transfused")), class = "data.frame", row.names = c(NA, 
-9L))

result <- transform_patient_data(df)
print(result)

运行后输出结果:

Patient_IDABC_ReturnedABC_TransfusedDEF_TransfusedGHI_Transfused
12344131

方法2:基础R实现(无依赖)

如果不想使用第三方包,可通过基础R函数完成转换:

transform_patient_data_base <- function(df) {
  # 生成三维交叉统计表
  tab <- table(df[c("Patient_ID", "Unit_Type", "Status")])
  # 转换为数据框格式
  tab_df <- as.data.frame(tab, responseName = "count")
  # 合并单元类型与状态为列名前缀
  tab_df$category <- paste(tab_df$Unit_Type, tab_df$Status, sep = "_")
  # 重塑为宽格式
  reshape(tab_df, 
          idvar = "Patient_ID", 
          timevar = "category", 
          direction = "wide",
          v.names = "count")
}

# 测试函数
result_base <- transform_patient_data_base(df)
print(result_base)

函数特性

  • 支持多患者ID批量处理,自动按患者单独统计
  • 缺失的单元类型-状态组合会自动填充0(tidyverse方法),避免NA值
  • 列名自动按Unit_Type_Status格式生成,完全匹配需求

内容的提问来源于stack exchange,提问作者Natasha H

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 18:55:40