Numpy中C/F连续数组sum与cumsum(axis=1)性能反常问题问询
为什么np.sum(axis=1)对Fortran-Contiguous数组更快,而np.cumsum(axis=1)却相反?
在测试不同存储顺序的NumPy数组性能时,我发现了一组矛盾的现象:
- 对
axis=1(行)求和时,Fortran-Contiguous(列优先)数组的性能反而优于C-Contiguous(行优先)数组,和“行优先数组行操作更缓存友好”的理论预期相反; - 但使用
np.cumsum(axis=1)时,C-Contiguous数组的性能符合预期,比Fortran-Contiguous数组快很多。
测试代码与结果
测试用数组定义:
import numpy as np arr_c = np.ones((1000, 1000), order='C') # 行优先存储 arr_f = np.ones((1000, 1000), order='F') # 列优先存储
np.sum(axis=1) 实测结果
%timeit arr_c.sum(axis=1) # 605 µs ± 40.8 µs per loop (mean ± std. dev. of 7 runs, 1,000 loops each) %timeit arr_f.sum(axis=1) # 398 µs ± 25.3 µs per loop (mean ± std. dev. of 7 runs, 1,000 loops each)
np.cumsum(axis=1) 实测结果
%timeit arr_c.cumsum(axis=1) # 4.01 ms ± 83 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) %timeit arr_f.cumsum(axis=1) # 15.6 ms ± 103 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)
现象原因分析
1. np.sum(axis=1):底层BLAS优化利用了Fortran-Contiguous的存储特性
NumPy的sum函数并非简单遍历累加,它会根据数组存储顺序和计算维度调用高度优化的BLAS库函数:
- 对于
axis=1的行求和,Fortran-Contiguous数组是列优先存储,此时行求和等价于对数组转置后做列求和(转置后的数组是C-Contiguous视图,无需复制内存)。 - BLAS库的列求和实现对连续内存块的处理效率极高——转置后的C-Contiguous数组中,列元素是连续存储的,缓存命中率拉满,性能自然更高。
- 而C-Contiguous数组的行求和,虽然行元素也是连续内存,但NumPy的sum实现未针对该场景做同等程度的BLAS优化,最终导致Fortran-Contiguous数组的求和速度反超。
2. np.cumsum(axis=1):必须按顺序遍历,缓存友好性起决定性作用
cumsum和sum的核心区别在于,它需要保留每一步的累加结果,无法像sum那样利用BLAS的批量优化,只能逐元素遍历计算:
- 对于C-Contiguous数组,
axis=1的行元素在内存中连续,遍历过程中缓存命中率极高,CPU可以持续从缓存读取数据,速度很快。 - 对于Fortran-Contiguous数组,
axis=1的行元素在内存中分散(每个元素间隔1000个内存单元),遍历会频繁触发缓存失效,需要从主存读取数据,导致速度大幅下降,完全符合缓存友好性的理论预期。
内容的提问来源于stack exchange,提问作者Kshitiz Sharma
相关产品推荐
相关产品推荐

