You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效获取data.table分组中的最后一条记录(无需max())

编辑:检查GForce问题

> sessionInfo()

R version 4.3.2 (2023-10-31 ucrt)

Platform: x86_64-w64-mingw32/x64 (64-bit)

Running under: Windows 10 x64 (build 19045)


Matrix products: default


locale:
[1] LC_COLLATE=German_Germany.utf8  LC_CTYPE=German_Germany.utf8    LC_MONETARY=German_Germany.utf8 LC_NUMERIC=C                   
[5] LC_TIME=German_Germany.utf8    


time zone: Europe/Berlin
tzcode source: internal


attached base packages:
[1] stats     graphics  grDevices utils     datasets  methods   base     


other attached packages:
[1] data.table_1.14.10


loaded via a namespace (and not attached):
[1] compiler_4.3.2 tools_4.3.2   

> Data <- data.table(id = c(rep("a", 2), rep("b",3)), 

+                    time = c(1:2, 1:3))

> Data[verbose = TRUE, , lastobs := max(time), by = id]

Detected that j uses these columns: time 

Finding groups using forderv ... forder.c received 5 rows and 1 columns

0.000s elapsed (0.000s cpu) 

Finding group sizes from the positions (can be avoided to save RAM) ... 
0.000s elapsed (0.000s cpu) 

lapply optimization is on, j unchanged as 'max(time)'

Old mean optimization is on, left j unchanged.

Making each group and running j (GForce FALSE) ... 
memcpy contiguous groups took 0.000s for 2 groups
eval(j) took 0.000s for 2 calls

0.000s elapsed (0.000s cpu) 

解决方案

方法1:直接取每组最后一行(依赖分组内顺序)

如果你的数据已经按time升序排列(每组内time从早到晚),用.N直接取每组最后一行是data.table里最高效的操作之一:

Data[, .SD[.N], by = id]

这里.SD指代每个分组的子数据集,.N是当前分组的总行数,.SD[.N]就精准定位到每组的最后一条记录。

方法2:先排序再取最后一行(确保取time最大的记录)

如果数据未按time排序,先通过setorder按id和time排序,再提取每组最后一行,速度远快于max()筛选:

# 先按id和time升序排序(time降序的话最后一行就是最大的,效果一样)
setorder(Data, id, time)
# 提取每组最后一行的索引并筛选
Data[Data[, .I[.N], by = id]$V1]

更简洁的链式写法:

setorder(Data, id, time)[, .SD[.N], by = id]

你的尝试出错原因

你写的Data[,.N == .I, by = id]中,.I是整个数据集的行号,而.N是当前分组的行数。第一个分组共2行,行号是1、2,只有行号2满足2 == .N;第二个分组共3行,行号是3、4、5,只有行号3满足3 == .N,这就是为什么会选中第一个分组最后一行和第二个分组第一行。

GForce优化补充

从你的verbose输出可见,max(time)未触发GForce优化(显示GForce FALSE),这也是原方案慢的原因。可以将年月变量转成整数类型(比如YYYYMM格式转成整数),触发GForce加速max()计算,但依然不如直接取分组最后一行高效:

# 假设time是Date/POSIXct类型,转成YYYYMM整数
Data[, time_num := as.integer(format(time, "%Y%m"))]
# 再运行原方案会触发GForce
Data[, lastobs := max(time_num), by = id]

内容的提问来源于stack exchange,提问作者Jakob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 12:23:14