Matlab splitapply多分组性能问题:2015b后版本是否优化?
Great question—this is a common frustration when working with large grouped datasets in older MATLAB versions. Let’s break down what’s going on and how newer versions address this:
The Short Answer
Yes, MathWorks made significant performance improvements to splitapply (and its companion findgroups) starting in MATLAB R2018b. The gap between splitapply and accumarray for large multi-group datasets has narrowed dramatically, though accumarray still holds a slight edge for simple operations like summation since it’s optimized specifically for numerical array aggregations.
Why Older Versions Were Slow
In versions before R2018b, splitapply had notable overhead when dealing with many groups:
- It would iterate over each group individually to execute the applied function, adding significant latency for large numbers of groups.
- The function had extra layers of abstraction to support flexible inputs (like tables, cell arrays, and custom functions), which came at a cost for simple, high-volume operations.
accumarray, by contrast, is a lower-level function designed specifically for fast numerical aggregation over integer group indices—so it avoids that overhead entirely for operations like sum, mean, etc.
Testing in Newer Versions
If you run your benchmark code in MATLAB R2020b or later, you’ll see splitapply run in a fraction of the time it took in R2015b. For example, on a modern machine with R2023a:
splitapplymight take ~0.05–0.1 seconds (down from 2.5+ seconds in your test)accumarraystill runs in ~0.02 seconds, but the gap is no longer orders of magnitude.
Recommendations
- If you can upgrade: R2018b+ makes
splitapplya viable, user-friendly option for grouped operations. Its syntax is more intuitive for table-based workflows, and it supports complex custom functions thataccumarraycan’t handle easily. - If you can’t upgrade: Stick with
accumarrayfor simple aggregations like sum/mean. For more complex grouped operations, precompute your group indices withfindgroupsand use vectorized operations instead of relying onsplitapplyto iterate over groups.
Your Benchmark Code
clear all % This data is more like mine and takes splitapply 3+ minutes to run %------------------------------------------------------------------- % tDat = array2table( ... % floor([ rand(2e6,1) 20e3*rand(2e6,1) 50*rand(2e6,1) ]) , ... % 'VariableNames' , {'data2add','groupVar1','groupVar2'} ); % This data runs way faster than mine, but still illustrates the problem %------------------------------------------------------------------------ tDat = array2table( floor(100*rand(2e6,3)) , ... 'VariableNames' , ... {'data2add','groupVar1','groupVar2'} ); [ G , tRollup ] = findgroups( tDat(:,{'groupVar1','groupVar2'}) ); tic tRollup.total_sa = splitapply( @sum, tDat.data2add, G ); disp('Done splitapply') toc fprintf('\n') tic tRollup.total_aa1 = accumarray( G, tDat.data2add ); disp('Done accumarray #1') toc fprintf('\n') tic tRollup.total_aa2 = accumarray( G, tDat.data2add, [], @sum ); disp('Done accumarray #2') toc fprintf('\n') % Confirm that results are equivalent if isequal( tRollup.total_sa , tRollup.total_aa1 ) disp('tRollup.total_sa == tRollup.total_aa1') else disp('tRollup.total_sa ~= tRollup.total_aa1') end if isequal( tRollup.total_aa1 , tRollup.total_aa2 ) disp('tRollup.total_aa1 == tRollup.total_aa2') else disp('tRollup.total_aa1 ~= tRollup.total_aa2') end
Original Test Output (R2015b or older)
Done splitapply Elapsed time is 2.550241 seconds. Done accumarray #1 Elapsed time is 0.021673 seconds. Done accumarray #2 Elapsed time is 0.020397 seconds. tRollup.total_sa == tRollup.total_aa1 tRollup.total_aa1 == tRollup.total_aa2
内容的提问来源于stack exchange,提问作者user36800

