R语言百万级数据下相邻行指定字段相等判断的高效实现
R语言高效实现相邻行多字段匹配判断
优化原理
原代码性能差是因为R的显式for循环属于上层解释执行,百万次迭代的索引、赋值开销极高,以下方案均采用底层C实现的向量化操作,性能提升可达数百倍。
方案1:基础R实现(无需安装额外包)
n <- nrow(data_ocurr_2019) # 第一行无前置行默认填NA,和原逻辑一致 data_ocurr_2019$Flag_Pac <- c(NA, as.integer(!( data_ocurr_2019$periodo_ocurr[-1] == data_ocurr_2019$periodo_ocurr[-n] & data_ocurr_2019$tipo_riesgo[-1] == data_ocurr_2019$tipo_riesgo[-n] & data_ocurr_2019$paciente[-1] == data_ocurr_2019$paciente[-n] )))
方案2:dplyr实现(可读性最优)
library(dplyr) data_ocurr_2019 <- data_ocurr_2019 %>% mutate(Flag_Pac = as.integer(!( periodo_ocurr == lag(periodo_ocurr) & tipo_riesgo == lag(tipo_riesgo) & paciente == lag(paciente) )))
方案3:data.table实现(百万级数据性能最优)
library(data.table) # 转换为data.table格式无内存拷贝,效率极高 setDT(data_ocurr_2019) data_ocurr_2019[, Flag_Pac := as.integer(!( periodo_ocurr == shift(periodo_ocurr) & tipo_riesgo == shift(tipo_riesgo) & paciente == shift(paciente) ))]
注:如果数据中存在NA值,
==比较会返回NA,可根据业务需求调整逻辑,比如将a == b替换为identical(a, b)或者%in%处理NA场景。
内容的提问来源于stack exchange,提问作者Carlos
相关产品推荐
相关产品推荐

