You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PAT_ID标记重复RESULT_TIME的data.table代码报错排查

按PAT_ID标记重复RESULT_TIME的问题排查与解决

需求说明

  • 新增列标记重复的RESULT_TIME,仅在同一PAT_ID分组内判断日期重复
  • 同一分组内重复日期的所有实例(包括首次出现)都要标记为重复

数据结构

c_petidesub_diab<-structure(list(PAT_ID = c(20509541L, 20996088L, 15699150L, 
9061383L, 
16105496L, 13781737L, 12814869L, 20723151L, 20451977L, 14145056L, 
14145056L, 14145056L, 19499987L, 19499987L, 19499987L, 19499987L, 
10206803L, 20935200L, 20205934L, 6471890L, 7312838L, 8705568L, 
12792099L, 19095082L, 19095082L, 19095082L, 20243600L, 20243600L, 
8285991L, 9175530L, 9175530L, 17169731L, 20793996L, 8892630L, 
8892630L, 12068391L, 5470224L, 21430152L, 6305585L, 6305585L), 
RESULT_TIME = structure(c(16386, 17200, 16931, 18500, 15453, 
16142, 17277, 17719, 17245, 16678, 16861, 18451, 15931, 15821, 
15861, 15793, 17634, 17141, 17769, 16659, 17611, 17964, 17709, 
16134, 16092, 16128, 15540, 18429, 16707, 15085, 15336, 18444, 
18548, 17452, 17452, 17961, 16149, 18121, 16408, 16427), class = "Date")), 
row.names = c(471L, 
1267L, 1880L, 3893L, 6410L, 7943L, 7979L, 9488L, 10163L, 11160L, 
11161L, 11162L, 11691L, 11692L, 11693L, 11694L, 11703L, 15464L, 
16526L, 17039L, 18436L, 19671L, 19946L, 20163L, 20164L, 20165L, 
20948L, 20949L, 22316L, 22411L, 22412L, 23885L, 24765L, 26629L, 
26630L, 26959L, 28229L, 32082L, 33923L, 33924L), class = "data.frame")

原代码及报错

原代码

c_peptide<-as.data.table(c_petidesub_diab)[, dup_c_peptide := 
(duplicated(c_petidesub_diab$RESULT_TIME))| 
(duplicated(c_petidesub_diab$RESULT_TIME, fromLast=TRUE)), by 
= c_petidesub_diab$PAT_ID][]

报错信息

Error in `[.data.table`(as.data.table(c_petidesub_diab), , `:=`(dup_c_peptide,  : 
Supplied 40 items to be assigned to group 1 of size 1 in column 'dup_c_peptide'. 
The RHS length must either be 1 (single values are ok) or match the LHS length 
exactly. If you wish to 'recycle' the RHS please use rep() explicitly to make this 
intent clear to readers of your code.

报错原因

  1. 分组参数错误:by = c_petidesub_diab$PAT_ID是传入整个原数据集的PAT_ID向量,会让data.table将每一行视为独立分组,而右侧计算的是全量数据集的重复标记(长度40),与单组长度1不匹配,触发报错。
  2. 重复判断范围错误:原代码直接引用c_petidesub_diab$RESULT_TIME,是基于全量数据集判断重复,没有按PAT_ID分组内的日期进行判断,不符合需求逻辑。

解决方案

使用data.table原生分组语法,直接引用列名作为分组键,并在分组内计算重复标记:

修正代码

library(data.table)
# 转换为data.table格式
setDT(c_petidesub_diab)
# 按PAT_ID分组,标记分组内重复的RESULT_TIME
c_petidesub_diab[, dup_c_peptide := duplicated(RESULT_TIME) | duplicated(RESULT_TIME, fromLast = TRUE), by = PAT_ID]

效果验证

以PAT_ID=8892630为例,两行的dup_c_peptide都会被标记为TRUE,完全符合需求。

补充:恢复数据框格式

如果需要将结果转回普通data.frame格式,可执行:

setDF(c_petidesub_diab)

内容的提问来源于stack exchange,提问作者stephr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 05:47:51