如何高效在R语言分组数据框中跨行匹配取值,无需使用for循环
解法
你可以直接用data.table的向量化操作实现,全程不需要for循环,性能远高于循环写法:
第一步:修正示例数据语法错误(原代码括号不规范会报错)
library(data.table) df <- data.table( personid = c(101, 102, 103, 104, 105, 201, 202, 203, 301, 302, 401), hh_id = c(1, 1, 1, 1, 1, 2, 2, 2, 3, 3, 4), fatherid = c(NA, NA, 101, 101, 101, NA, NA, 201, NA, NA, NA), cancer = c(1,0,0,0,0,0,0,0,0,0,1) )
第二步:匹配父亲患癌状态
提供两种常用实现方案:
- 小数据量首选:用
match匹配,写法极简
df[, fathercancer := cancer[match(fatherid, personid)]]
- 大数据量首选:用
data.table自连接,性能更高
df[df, on = .(personid = fatherid), fathercancer := i.cancer]
运行后得到的fathercancer列和你给出的预期结果完全一致。
内容的提问来源于stack exchange,提问作者David Gil Solsona
相关产品推荐
相关产品推荐

