R语言优化循环效率:基于另一数据框行操作目标数据框列
跨数据框列运算的高效实现方案
数据准备
我有两个数据框:
dat <- data.frame(Digits_Lower = 1:5, Digits_Upper = 6:10, random = 20:24) dat #> Digits_Lower Digits_Upper random #> 1 1 6 20 #> 2 2 7 21 #> 3 3 8 22 #> 4 4 9 23 #> 5 5 10 24 cb <- data.frame(Digits = c("Digits_Lower", "Digits_Upper"), x = 1:2, y = 3:4) cb #> Digits x y #> 1 Digits_Lower 1 3 #> 2 Digits_Upper 2 4
需求与现有循环实现
我需要对dat中的指定列执行运算:将dat的列与cb中对应行按公式(dat列值 - cb$y[i]) * cb$x[i]计算,生成新列。目前用for循环实现如下:
dat.loop <- dat for(i in seq_len(nrow(cb))) { # 基于cb的Digits列创建新列 dat.loop[paste0("disp", sep = '.', cb$Digits[i])] <- # 将dat列中每个值与cb对应行执行运算 (dat.loop[, cb$Digits[i]]- cb$y[i]) * cb$x[i] } dat.loop #> Digits_Lower Digits_Upper random disp.Digits_Lower disp.Digits_Upper #> 1 1 6 20 -2 4 #> 2 2 7 21 -1 6 #> 3 3 8 22 0 8 #> 4 4 9 23 1 10 #> 5 5 10 24 2 12
但我的数据集规模极大,循环使用起来非常繁琐,后续还要对新增列执行类似运算。
疑问
- 使用
data.table或tidyverse这类工具能否提升运算效率? - 如何优化现有循环或用其他方法编写简洁代码,实现
dat列与cb对应行的运算?理想状态是把循环拆成多个函数,避免重复索引cb或添加不必要的数据,但不知道怎么实现。
编辑补充:通用化实现(基于tidytable)
我修改了他人提供的代码,使其更通用——因为dat和cb来自用户上传文件,dat的列数和列名不固定,但需要处理的列均以"Digits_"开头,且在cb的Digits列中指定:
library(tidytable) results <- dat %>% crossing.(cb) %>% mutate_rowwise.(disp = (get(`Digits`)-y) *x ) %>% pivot_wider.(names_from = Digits, values_from = disp, names_prefix = "disp_") results2 <- results %>% fill.(starts_with("disp"), .direction = c("downup"), .by = 'random') %>% select.(-c(x,y)) %>% distinct.() results2 #> Digits_Lower Digits_Upper random disp_Digits_Lower disp_Digits_Upper #> 1 1 6 20 -2 4 #> 2 2 7 21 -1 6 #> 3 3 8 22 0 8 #> 4 4 9 23 1 10 #> 5 5 10 24 2 12
内容的提问来源于stack exchange,提问作者kzyat
相关产品推荐
相关产品推荐

