基于PAT_ID标记重复RESULT_TIME的data.table代码报错排查
按PAT_ID标记重复RESULT_TIME的问题排查与解决
需求说明
- 新增列标记重复的
RESULT_TIME,仅在同一PAT_ID分组内判断日期重复 - 同一分组内重复日期的所有实例(包括首次出现)都要标记为重复
数据结构
c_petidesub_diab<-structure(list(PAT_ID = c(20509541L, 20996088L, 15699150L, 9061383L, 16105496L, 13781737L, 12814869L, 20723151L, 20451977L, 14145056L, 14145056L, 14145056L, 19499987L, 19499987L, 19499987L, 19499987L, 10206803L, 20935200L, 20205934L, 6471890L, 7312838L, 8705568L, 12792099L, 19095082L, 19095082L, 19095082L, 20243600L, 20243600L, 8285991L, 9175530L, 9175530L, 17169731L, 20793996L, 8892630L, 8892630L, 12068391L, 5470224L, 21430152L, 6305585L, 6305585L), RESULT_TIME = structure(c(16386, 17200, 16931, 18500, 15453, 16142, 17277, 17719, 17245, 16678, 16861, 18451, 15931, 15821, 15861, 15793, 17634, 17141, 17769, 16659, 17611, 17964, 17709, 16134, 16092, 16128, 15540, 18429, 16707, 15085, 15336, 18444, 18548, 17452, 17452, 17961, 16149, 18121, 16408, 16427), class = "Date")), row.names = c(471L, 1267L, 1880L, 3893L, 6410L, 7943L, 7979L, 9488L, 10163L, 11160L, 11161L, 11162L, 11691L, 11692L, 11693L, 11694L, 11703L, 15464L, 16526L, 17039L, 18436L, 19671L, 19946L, 20163L, 20164L, 20165L, 20948L, 20949L, 22316L, 22411L, 22412L, 23885L, 24765L, 26629L, 26630L, 26959L, 28229L, 32082L, 33923L, 33924L), class = "data.frame")
原代码及报错
原代码
c_peptide<-as.data.table(c_petidesub_diab)[, dup_c_peptide := (duplicated(c_petidesub_diab$RESULT_TIME))| (duplicated(c_petidesub_diab$RESULT_TIME, fromLast=TRUE)), by = c_petidesub_diab$PAT_ID][]
报错信息
Error in `[.data.table`(as.data.table(c_petidesub_diab), , `:=`(dup_c_peptide, : Supplied 40 items to be assigned to group 1 of size 1 in column 'dup_c_peptide'. The RHS length must either be 1 (single values are ok) or match the LHS length exactly. If you wish to 'recycle' the RHS please use rep() explicitly to make this intent clear to readers of your code.
报错原因
- 分组参数错误:
by = c_petidesub_diab$PAT_ID是传入整个原数据集的PAT_ID向量,会让data.table将每一行视为独立分组,而右侧计算的是全量数据集的重复标记(长度40),与单组长度1不匹配,触发报错。 - 重复判断范围错误:原代码直接引用
c_petidesub_diab$RESULT_TIME,是基于全量数据集判断重复,没有按PAT_ID分组内的日期进行判断,不符合需求逻辑。
解决方案
使用data.table原生分组语法,直接引用列名作为分组键,并在分组内计算重复标记:
修正代码
library(data.table) # 转换为data.table格式 setDT(c_petidesub_diab) # 按PAT_ID分组,标记分组内重复的RESULT_TIME c_petidesub_diab[, dup_c_peptide := duplicated(RESULT_TIME) | duplicated(RESULT_TIME, fromLast = TRUE), by = PAT_ID]
效果验证
以PAT_ID=8892630为例,两行的dup_c_peptide都会被标记为TRUE,完全符合需求。
补充:恢复数据框格式
如果需要将结果转回普通data.frame格式,可执行:
setDF(c_petidesub_diab)
内容的提问来源于stack exchange,提问作者stephr
相关产品推荐
相关产品推荐

