在基础R中,先选列再过滤行为何比先过滤行再选列更快?
为什么
df$type[df$weight < Treshold] <- "M"比df[df$weight < Treshold, ]$type <- "M"更快? 先看这段用于根据weight列值修改type列的R代码,以及对应的性能测试:
n <- 1e3; m <- n*10 Treshold <- 50 wts <- runif(m) df <- data.frame(id=seq_len(m), weight=wts * 100, type='L') library(microbenchmark) microbenchmark( "df-col-row" = (df$type[df$weight < Treshold] <- "M"), "df-row-col" = (df[df$weight < Treshold, ]$type <- "M") )
性能测试结果:
Unit: microseconds
expr min lq mean median uq max neval
df-col-row 80.6 87.65 145.429 89.55 104.55 5109.1 100
df-row-col 564.9 586.10 618.496 592.40 618.90 1601.0 100
两种写法的性能差距明显,第一种方式快很多,原因是什么?
更新1:列数越多,性能差异越大
当数据框包含更多列时,第二种写法的性能下降更明显:
d9 <- data.frame(type='L', weight=wts * 100, c3=3, c4=4, c5=5, c6=6, c7=7, c8=8, c9=9) microbenchmark( "df-row-9col" = (d9[d9$weight < Treshold, ]$type <- "M") )
测试结果:
Unit: microseconds
expr min lq mean median uq max neval
df-row-9col 950.1 1091.55 1267.982 1111.1 1172.45 5806 100
更新2:内存复制次数差异
用tracemem追踪内存变化可以看到两种写法的核心差异:
tracemem(df) df$type[df$weight < Treshold] <- "M" # 方式1 # tracemem[0x000002c92d2b87c8 -> 0x000002c92d2b9498]: $<-.data.frame $<- df[df$weight < Treshold, ]$type <- "M" # 方式2 # tracemem[0x000002c92d2b9498 -> 0x000002c92d2b9ad8]: # tracemem[0x000002c92d2b9ad8 -> 0x000002c92d2c47d8]: [<-.[<- untracemem(df)
原因解释
- 方式1:直接定位到
type列,再筛选需要修改的行,整个过程仅触发一次内存复制,修改操作直接作用在目标列的指定位置,开销极小。 - 方式2:先通过
df[df$weight < Treshold, ]筛选符合条件的行,这一步会**复制整个数据框的对应行(包含所有列)**生成临时数据框;之后修改临时数据框的type列,最后还要将修改内容同步回原数据框,全程触发两次内存复制。当数据框列数越多,第一次复制的临时数据体积越大,性能开销越高,这也解释了更新1中列数增加后性能差异进一步扩大的现象。
内容的提问来源于stack exchange,提问作者clp
相关产品推荐
相关产品推荐

