You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于优先级列表筛选R data.table中的诊断值

按优先级保留单一诊断的data.table解决方案

思路

核心逻辑:先定位每个患者优先级最高的诊断(按diagnosis1→diagnosis2→diagnosis3→diagnosis4的顺序,取第一个出现1的诊断列),再将所有诊断列重置为0,最后仅保留该最高优先级诊断的1值,同时保留其他原有变量。

代码实现

library(data.table)

# 示例数据集
dt <- data.table(ID = c(1,2,3, 4, 5),
                 diagnosis1 = c(0, 0, 1, 0, 1), 
                 diagnosis2 = c(1, 0, 0, 1, 0), 
                 diagnosis3 = c(0, 1, 1, 0, 1), 
                 diagnosis4 = c(1, 0, 1, 0, 0))

# 定义诊断列的优先级顺序
diagnosis_cols <- c("diagnosis1", "diagnosis2", "diagnosis3", "diagnosis4")

# 为每个患者确定最高优先级的诊断(第一个值为1的诊断列)
dt[, highest_diag := {
  # 找到当前行中第一个值为1的诊断列索引
  first_pos <- which(.SD == 1)[1]
  # 存在有效诊断则返回对应列名,否则返回NA
  if (length(first_pos) > 0) diagnosis_cols[first_pos] else NA
}, .SDcols = diagnosis_cols, by = ID]

# 将所有诊断列初始化为0
dt[, (diagnosis_cols) := 0]

# 仅将最高优先级诊断列设为1(针对有诊断的患者)
dt[!is.na(highest_diag), (highest_diag) := 1, by = ID]

# (可选)删除临时辅助列
dt[, highest_diag := NULL]

结果验证

处理后的数据集如下:

ID diagnosis1 diagnosis2 diagnosis3 diagnosis4
1:  1          0          1          0          0
2:  2          0          0          1          0
3:  3          1          0          0          0
4:  4          0          1          0          0
5:  5          1          0          0          0

每个患者仅保留优先级最高的诊断,其他诊断列均为0,同时保留了ID等原有变量,完全符合需求。

内容的提问来源于stack exchange,提问作者Hellihansen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 05:12:02