You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在基础R中,先选列再过滤行为何比先过滤行再选列更快?

为什么df$type[df$weight < Treshold] <- "M"比df[df$weight < Treshold, ]$type <- "M"更快?

先看这段用于根据weight列值修改type列的R代码,以及对应的性能测试:

n <- 1e3; m <- n*10
Treshold <- 50
wts      <- runif(m)
df  <- data.frame(id=seq_len(m), weight=wts * 100, type='L')

library(microbenchmark)
microbenchmark(
"df-col-row" = (df$type[df$weight < Treshold]   <- "M"),
"df-row-col" = (df[df$weight < Treshold, ]$type <- "M")
)

性能测试结果:

Unit: microseconds
expr min lq mean median uq max neval
df-col-row 80.6 87.65 145.429 89.55 104.55 5109.1 100
df-row-col 564.9 586.10 618.496 592.40 618.90 1601.0 100

两种写法的性能差距明显,第一种方式快很多,原因是什么?


更新1:列数越多,性能差异越大

当数据框包含更多列时,第二种写法的性能下降更明显:

d9  <- data.frame(type='L', weight=wts * 100, c3=3, c4=4, c5=5, c6=6, c7=7, c8=8, c9=9)
microbenchmark(
"df-row-9col" = (d9[d9$weight < Treshold, ]$type <- "M")
)

测试结果:

Unit: microseconds
expr min lq mean median uq max neval
df-row-9col 950.1 1091.55 1267.982 1111.1 1172.45 5806 100


更新2:内存复制次数差异

用tracemem追踪内存变化可以看到两种写法的核心差异:

tracemem(df)
df$type[df$weight < Treshold]   <- "M"    # 方式1
# tracemem[0x000002c92d2b87c8 -> 0x000002c92d2b9498]: $<-.data.frame $<- 

df[df$weight < Treshold, ]$type <- "M"    # 方式2
# tracemem[0x000002c92d2b9498 -> 0x000002c92d2b9ad8]: 
# tracemem[0x000002c92d2b9ad8 -> 0x000002c92d2c47d8]: [<-.[<-
untracemem(df)

原因解释

  • 方式1:直接定位到type列,再筛选需要修改的行,整个过程仅触发一次内存复制,修改操作直接作用在目标列的指定位置,开销极小。
  • 方式2:先通过df[df$weight < Treshold, ]筛选符合条件的行,这一步会**复制整个数据框的对应行(包含所有列)**生成临时数据框;之后修改临时数据框的type列,最后还要将修改内容同步回原数据框,全程触发两次内存复制。当数据框列数越多,第一次复制的临时数据体积越大,性能开销越高,这也解释了更新1中列数增加后性能差异进一步扩大的现象。

内容的提问来源于stack exchange,提问作者clp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 19:22:54