You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按患者保留首个保险ID过滤数据集的实现代码咨询

R实现:保留每位患者最早保险编号对应的全部记录

需求说明

现有患者保险编号列表,部分患者被错误分配了多个保单号,需仅保留每位患者最早出现的保险编号对应的所有记录,删除同一患者其他保险编号的记录。

测试数据构造

InsuranceNumber <- c("00932", "00932", "00932", "00987", "00987", "00915", "00915", "00923" , "00977")
PatientName <- c("Patient1", "Patient1", "Patient1", "Patient1", "Patient1", "Patient1", "Patient1", "Patient2", "Patient2")
df <- data.frame(InsuranceNumber, PatientName)

实现代码

方法1:dplyr包实现(推荐,语法简洁)

library(dplyr)
df_result <- df %>%
  group_by(PatientName) %>%
  filter(InsuranceNumber == first(InsuranceNumber)) %>%
  ungroup()

方法2:基础R实现(无需额外安装依赖包)

# 提取每个患者首次出现的保险编号
first_ins <- aggregate(InsuranceNumber ~ PatientName, df, FUN = function(x) x[1])
# 匹配保留对应记录
df_result <- merge(df, first_ins, by = c("PatientName", "InsuranceNumber"))

结果验证

运行print(df_result)即可得到预期输出:

InsuranceNumber PatientName
1           00932    Patient1
2           00932    Patient1
3           00932    Patient1
4           00923    Patient2

内容的提问来源于stack exchange,提问作者Random Person

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 01:15:03