You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何NumPy ufuncs(如np.cumsum)在某轴上运算速度是另一轴的2倍?

Why is np.cumsum(axis=1) nearly twice as fast as axis=0?

Great question! Let's start by recapping the performance test you ran to set the context:

import numpy as np

# Create a 1000x1000 array (total 1 million elements)
arr = np.arange(int(1E6)).reshape(int(1E3), -1)

In [52]: %timeit arr.cumsum(axis=1)
2.27 ms ± 10.5 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)

In [53]: %timeit arr.cumsum(axis=0)
4.16 ms ± 10.3 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)

As you noticed, row-wise cumulative sum (along axis=1) is almost twice as fast as column-wise (along axis=0). The root cause boils down to how NumPy stores data in memory and how CPUs handle memory access. Let's break it down step by step:


1. Memory Layout & Cache Locality is King

By default, NumPy arrays use row-major (C-style) memory ordering. That means every element in a row is stored right next to each other in memory, while elements in the same column are spread out with a big "jump" (stride) between them—1000 elements apart in your 1000x1000 array.

  • When you run cumsum(axis=1), the operation reads memory in contiguous chunks. CPUs rely heavily on cache memory (much faster than main RAM), and contiguous access lets the cache load entire blocks of data at once. This minimizes "cache misses" (times the CPU has to wait for data from main RAM), making the operation super efficient.
  • When you run cumsum(axis=0), you're skipping 1000 elements each time to grab the next value in the column. This scattered access means the cache can't keep up—most of the time, the data you need isn't in the cache, forcing the CPU to wait for slow main RAM access. That's the biggest hit to performance here.

2. Vectorization Works Best on Contiguous Data

NumPy's ufuncs like cumsum are optimized to use CPU vector instructions (SIMD) that process multiple elements at once. But this only works smoothly when data is contiguous:

  • For row-wise operations, the contiguous block of data fits perfectly into the CPU's vector registers, letting it crunch through elements in parallel.
  • For column-wise operations, the non-contiguous memory access breaks this parallel flow. The CPU has to fetch small, scattered chunks of data instead of processing a big block, which wastes cycles and slows things down.

3. Optimized Implementations Favor Contiguous Axes

Under the hood, NumPy's cumsum uses highly optimized C code or BLAS routines. These implementations are written to take full advantage of contiguous memory layouts. When you work on a non-contiguous axis, the code has to handle extra memory addressing logic, adding overhead that you don't see with the contiguous case.


Want to test this yourself? Try creating a column-major (Fortran-style) array with np.arange(int(1E6)).reshape(int(1E3), -1, order='F')—you'll see that cumsum(axis=0) becomes faster than axis=1 because the memory layout now favors column-wise access.

内容的提问来源于stack exchange,提问作者kmario23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:38:06