You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何numpy数组在multiprocessing库中处理速度远慢于列表?

多进程处理Numpy数组性能异常问题

设备与环境

  • 处理器:两台AMD 7302 16核处理器(总计32核)
  • 系统:Red Hat 8.4
  • Python版本:3.10.6

测试代码

为学习multiprocessing库编写的测试代码如下:

from multiprocessing import Pool
import numpy as np
import sys
import datetime

def f(x):
    return x**2

def main(DataType="List", NThr=2, Vectorize=False):
    N = 5*10**7           # number of elements
    n = NThr              # number of threads
    y = np.zeros(N)
    # Use list
    if(DataType == "List"):
        x = []
        for i in range(N):
            x.append(i)
    # Use Numpy
    elif(DataType=="Numpy"):
        x = np.zeros(N)
        for i in range(len(x)):
            x[i] = i
    # Run parallel code
    t0 = datetime.datetime.now()
    if(n==1):
        if(DataType == "Numpy" and Vectorize == True):
            y = np.vectorize(f)(x)
        else:
            for i in range(len(x)):
                y[i] = f(x[i])
    else:
        with Pool(n) as p:
            y = p.map(f, x)
    t1 = datetime.datetime.now()
    dt = (t1 - t0).total_seconds()
    print("{} : Vect = {}, n = {}, time : {}s".format(DataType,Vectorize,n,dt))
    sys.exit(0)

if __name__ == "__main__":
    main()

测试结果

多次运行后得到以下耗时数据:

Numpy : Vect = True, n = 1, time : 9.566441s
Numpy : Vect = False, n = 1, time : 16.00333s
Numpy : Vect = False, n = 2, time : 143.331352s
List : Vect = False, n = 1, time : 21.11657s
List : Vect = False, n = 2, time : 11.868897s
List : Vect = False, n = 5, time : 6.162561s

其中Numpy数组+2进程的耗时(143秒)远高于列表+2进程(11.9秒),甚至比单进程处理Numpy数组慢很多。

性能分析结果

使用cProfile对两个版本进行性能分析,结果如下:

Numpy版本性能分析

ncalls  tottime  percall  cumtime  percall filename:lineno(function)
# Time consuming
1    0.000    0.000  138.997  138.997 pool.py:362(map)
1    0.000    0.000  138.956  138.956 pool.py:764(wait)
1    0.000    0.000  138.956  138.956 pool.py:767(get)
4    0.000    0.000  138.957   34.739 threading.py:288(wait)
4    0.000    0.000  138.957   34.739 threading.py:589(wait)
14/1    0.000    0.000  145.150  145.150 {built-in method builtins.exec}
19  138.957    7.314  138.957    7.314 {method 'acquire' of '_thread.lock' objects}
# Different number of calls
6    0.000    0.000    0.088    0.015 popen_fork.py:24(poll)
1    0.000    0.000    0.088    0.088 popen_fork.py:36(wait)
1    0.000    0.000    0.088    0.088 process.py:142(join)
10    0.000    0.000    0.000    0.000 process.py:99(_check_closed)
18    0.000    0.000    0.000    0.000 util.py:48(debug)
76    0.000    0.000    0.000    0.000 {built-in method builtins.len}
2    0.000    0.000    0.000    0.000 {built-in method numpy.zeros}
17    0.000    0.000    0.000    0.000 {built-in method posix.getpid}
6    0.088    0.015    0.088    0.015 {built-in method posix.waitpid}
3    0.000    0.000    0.000    0.000 {method 'append' of 'list' objects}

List版本性能分析

ncalls  tottime  percall  cumtime  percall filename:lineno(function)
# Time consuming
1    0.000    0.000   13.961   13.961 pool.py:362(map)
1    0.000    0.000   13.920   13.920 pool.py:764(wait)
1    0.000    0.000   13.920   13.920 pool.py:767(get)
4    0.000    0.000   13.921    3.480 threading.py:288(wait)
4    0.000    0.000   13.921    3.480 threading.py:589(wait)
14/1    0.000    0.000   24.475   24.475 {built-in method builtins.exec}
19   13.921    0.733   13.921    0.733 {method 'acquire' of '_thread.lock' objects}
# Different number of calls
7    0.000    0.000    0.132    0.019 popen_fork.py:24(poll)
2    0.000    0.000    0.132    0.066 popen_fork.py:36(wait)
2    0.000    0.000    0.132    0.066 process.py:142(join)
12    0.000    0.000    0.000    0.000 process.py:99(_check_closed)
19    0.000    0.000    0.000    0.000 util.py:48(debug)
75    0.000    0.000    0.000    0.000 {built-in method builtins.len}
1    0.000    0.000    0.000    0.000 {built-in method numpy.zeros}
18    0.000    0.000    0.000    0.000 {built-in method posix.getpid}
7    0.132    0.019    0.132    0.019 {built-in method posix.waitpid}
50000003    2.780    0.000    2.780    0.000 {method 'append' of 'list' objects}

注:List版本初始化x时调用了50000003次append(),而Numpy版本仅调用3次。

核心问题

为何numpy数组在multiprocessing库中运行耗时如此之长,尤其是当NThr==2时?


原因分析

  1. 进程间数据传递的开销差异

    • 处理列表时,Pool.map()会将列表分割为若干块后通过进程间通信(IPC)传递给子进程,列表元素是Python原生int对象,分割和传递开销可控。
    • 处理Numpy数组时,Pool.map()会遍历数组的每个元素,逐个通过pickle序列化传递——5亿个元素的逐个序列化/反序列化操作产生了巨量IPC开销,这是性能暴跌的核心原因。此外,Numpy数组作为连续内存块,默认传递时会触发完整内存拷贝,进一步放大开销。
  2. 违背Numpy的设计优化方向
    Numpy的核心优势是单进程下的批量向量运算,np.vectorize()虽模拟向量化,但本质仍是基于批量处理的优化;而用multiprocessing逐元素处理Numpy数组,完全放弃了其批量处理优势,反而暴露了进程间通信的劣势。

  3. 锁竞争的量级放大
    从性能分析结果可见,Numpy版本的锁acquire操作耗时138秒,远高于List版本的13秒。这是因为Numpy版本的进程间通信次数(5亿次)远多于List版本(分块传递,单次传递元素数量多),导致进程间锁竞争和IPC等待次数呈数量级增长,进一步拖慢整体速度。


内容的提问来源于stack exchange,提问作者irritable_phd_syndrome

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 02:46:18