You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

150万行data.table逐行调用自定义函数效率低的原因排查

Troubleshooting Slow Row-wise Function Application on Large data.table

Hey there, let's break this down because dealing with 1.5 million rows in data.table can get really sluggish if you're not using its optimized workflows properly. First off—using by=1:nrow(dt) is almost certainly a huge part of your performance problem, but we'll also cover how to check if your growth.ls function is contributing to the slowdown.

Why by=1:nrow(dt) is a Performance Killer

data.table's by argument is designed for grouping rows into meaningful, larger subsets (e.g., by a category column) for aggregation. When you use by=1:nrow(dt), you're forcing it to create 1.5 million separate groups, each containing exactly one row. This introduces massive overhead: memory allocation for each tiny group, repeated context switching, and redundant setup/teardown processes that completely negate data.table's speed advantages.

Step 1: Fix the Calling Pattern First

Let's replace that inefficient row-wise grouping with approaches that leverage data.table's vectorized strengths:

Option 1: Vectorize growth.ls (Best Practice)

If you can modify growth.ls to accept entire columns (vectors) instead of single-row values, you can run it once on the entire dataset in a single pass. For example:

# Assuming growth.ls is updated to handle column vectors
dt[, result := growth.ls(.SD), .SDcols = scols]

This is the fastest possible approach because vectorized operations in R avoid per-row loop overhead entirely.

Option 2: Use Parallelized Row-wise Mapping (If Vectorization Isn't Possible)

If growth.ls must operate on individual rows (e.g., it relies on row-specific context or non-vectorizable logic), use purrr::pmap instead of by=1:nrow(dt). It avoids the grouping overhead of data.table's by while handling row-wise inputs cleanly:

library(purrr)
# pmap passes each row's scols values as arguments to growth.ls
dt[, result := pmap(.SD, growth.ls), .SDcols = scols]

For even faster results, pair this with furrr to parallelize across CPU cores:

library(furrr)
plan(multisession) # Use all available cores
dt[, result := future_pmap(.SD, growth.ls), .SDcols = scols]

Option 3: Rewrite growth.ls in Rcpp (For Extreme Speed)

If growth.ls is a pure R function with internal loops, rewriting it in C++ via Rcpp can yield 10-100x speedups. data.table plays nicely with Rcpp functions, so you can call your compiled function directly on the columns.

Step 2: Diagnose if growth.ls Is the Bottleneck

Even with a better calling pattern, a poorly optimized growth.ls can still slow things down. Here's how to check:

  • Test with a small sample: Grab 1,000 rows and compare the time of vectorized vs. row-wise calls using microbenchmark:
    library(microbenchmark)
    sample_dt <- dt[1:1000, sccols, with = FALSE]
    
    bench_results <- microbenchmark(
      vectorized = growth.ls(sample_dt),
      rowwise = lapply(1:nrow(sample_dt), function(i) growth.ls(sample_dt[i,])),
      times = 10
    )
    print(bench_results)
    
    If the vectorized version is drastically faster, your function can be optimized for vectorization.
  • Profile the code: Use profvis to see exactly where time is being spent. Run this on a small subset to avoid waiting forever:
    library(profvis)
    profvis({
      dt[1:10000, growth.ls(.SD), by=1:nrow(dt), .SDcols=scols]
    })
    
    The flame graph will show if time is going to data.table's grouping logic or your function's internal computations.

Quick Summary

9 times out of 10, the by=1:nrow(dt) pattern is the main culprit here. Switching to vectorized operations or pmap should give you a massive speed boost. If the function itself is still slow, use profiling to target optimizations like removing internal loops or switching to Rcpp.

内容的提问来源于stack exchange,提问作者Mohamed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:12:09