150万行data.table逐行调用自定义函数效率低的原因排查
Hey there, let's break this down because dealing with 1.5 million rows in data.table can get really sluggish if you're not using its optimized workflows properly. First off—using by=1:nrow(dt) is almost certainly a huge part of your performance problem, but we'll also cover how to check if your growth.ls function is contributing to the slowdown.
Why by=1:nrow(dt) is a Performance Killer
data.table's by argument is designed for grouping rows into meaningful, larger subsets (e.g., by a category column) for aggregation. When you use by=1:nrow(dt), you're forcing it to create 1.5 million separate groups, each containing exactly one row. This introduces massive overhead: memory allocation for each tiny group, repeated context switching, and redundant setup/teardown processes that completely negate data.table's speed advantages.
Step 1: Fix the Calling Pattern First
Let's replace that inefficient row-wise grouping with approaches that leverage data.table's vectorized strengths:
Option 1: Vectorize growth.ls (Best Practice)
If you can modify growth.ls to accept entire columns (vectors) instead of single-row values, you can run it once on the entire dataset in a single pass. For example:
# Assuming growth.ls is updated to handle column vectors dt[, result := growth.ls(.SD), .SDcols = scols]
This is the fastest possible approach because vectorized operations in R avoid per-row loop overhead entirely.
Option 2: Use Parallelized Row-wise Mapping (If Vectorization Isn't Possible)
If growth.ls must operate on individual rows (e.g., it relies on row-specific context or non-vectorizable logic), use purrr::pmap instead of by=1:nrow(dt). It avoids the grouping overhead of data.table's by while handling row-wise inputs cleanly:
library(purrr) # pmap passes each row's scols values as arguments to growth.ls dt[, result := pmap(.SD, growth.ls), .SDcols = scols]
For even faster results, pair this with furrr to parallelize across CPU cores:
library(furrr) plan(multisession) # Use all available cores dt[, result := future_pmap(.SD, growth.ls), .SDcols = scols]
Option 3: Rewrite growth.ls in Rcpp (For Extreme Speed)
If growth.ls is a pure R function with internal loops, rewriting it in C++ via Rcpp can yield 10-100x speedups. data.table plays nicely with Rcpp functions, so you can call your compiled function directly on the columns.
Step 2: Diagnose if growth.ls Is the Bottleneck
Even with a better calling pattern, a poorly optimized growth.ls can still slow things down. Here's how to check:
- Test with a small sample: Grab 1,000 rows and compare the time of vectorized vs. row-wise calls using
microbenchmark:
If the vectorized version is drastically faster, your function can be optimized for vectorization.library(microbenchmark) sample_dt <- dt[1:1000, sccols, with = FALSE] bench_results <- microbenchmark( vectorized = growth.ls(sample_dt), rowwise = lapply(1:nrow(sample_dt), function(i) growth.ls(sample_dt[i,])), times = 10 ) print(bench_results) - Profile the code: Use
profvisto see exactly where time is being spent. Run this on a small subset to avoid waiting forever:
The flame graph will show if time is going tolibrary(profvis) profvis({ dt[1:10000, growth.ls(.SD), by=1:nrow(dt), .SDcols=scols] })data.table's grouping logic or your function's internal computations.
Quick Summary
9 times out of 10, the by=1:nrow(dt) pattern is the main culprit here. Switching to vectorized operations or pmap should give you a massive speed boost. If the function itself is still slow, use profiling to target optimizations like removing internal loops or switching to Rcpp.
内容的提问来源于stack exchange,提问作者Mohamed

