是否有更快速的方法对数据框指定列按行求和,替代dplyr的rowwise方法?
性能优化方案
原代码的性能瓶颈来自rowwise()逐行迭代的高额开销,你的计算逻辑完全可以用向量化操作替代,无需逐行处理。
方案1:适配当前场景的最简最高效写法
你要计算的channel.corr本质是ladder.channel指定列与correction列的数值相加,直接用+运算符即可,完全不需要调用sum()和rowwise():
library(dplyr) x <- data.frame("channel_1" = seq(1, 10), "channel_2" = seq(1, 10), "channel_3" = seq(1, 10), "channel_4" = seq(1, 10), "channel_5" = seq(1, 10)) ladder.channel <- "channel_4" bleed.channels <- setdiff(c("channel_1", "channel_2", "channel_3", "channel_4", "channel_5"), ladder.channel) y <- x %>% mutate( correction = -pmax(!!!syms(bleed.channels)), channel.corr = .data[[ladder.channel]] + correction )
方案2:通用多列按行求和方案
如果后续需要对更多字符向量指定的列按行求和,使用rowSums()替代逐行sum,性能同样远高于rowwise:
y <- x %>% mutate( correction = -pmax(!!!syms(bleed.channels)), channel.corr = rowSums(across(all_of(c(ladder.channel, "correction")))) )
性能差异说明
- 向量化运算(
+/rowSums)底层为C语言实现的批量运算,百万行级数据计算耗时通常低于1秒 rowwise()为R层面的逐行迭代逻辑,每行都会产生额外的调用开销,数据量越大性能衰减越严重
内容的提问来源于stack exchange,提问作者Mike
相关产品推荐
相关产品推荐

