R语言按患者保留首个保险ID过滤数据集的实现代码咨询
R实现:保留每位患者最早保险编号对应的全部记录
需求说明
现有患者保险编号列表,部分患者被错误分配了多个保单号,需仅保留每位患者最早出现的保险编号对应的所有记录,删除同一患者其他保险编号的记录。
测试数据构造
InsuranceNumber <- c("00932", "00932", "00932", "00987", "00987", "00915", "00915", "00923" , "00977") PatientName <- c("Patient1", "Patient1", "Patient1", "Patient1", "Patient1", "Patient1", "Patient1", "Patient2", "Patient2") df <- data.frame(InsuranceNumber, PatientName)
实现代码
方法1:dplyr包实现(推荐,语法简洁)
library(dplyr) df_result <- df %>% group_by(PatientName) %>% filter(InsuranceNumber == first(InsuranceNumber)) %>% ungroup()
方法2:基础R实现(无需额外安装依赖包)
# 提取每个患者首次出现的保险编号 first_ins <- aggregate(InsuranceNumber ~ PatientName, df, FUN = function(x) x[1]) # 匹配保留对应记录 df_result <- merge(df, first_ins, by = c("PatientName", "InsuranceNumber"))
结果验证
运行print(df_result)即可得到预期输出:
InsuranceNumber PatientName 1 00932 Patient1 2 00932 Patient1 3 00932 Patient1 4 00923 Patient2
内容的提问来源于stack exchange,提问作者Random Person
相关产品推荐
相关产品推荐

