You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否用multiprocessing加速大尺寸numpy数组的np.sum计算?

问题

当处理尺寸在10⁸到10⁹之间的numpy数组时,是否存在比np.sum更快的求和方式?我尝试使用fork模式的multiprocessing进行并行求和,但无论设置1-4个worker,其速度都慢于直接调用np.sum。当前环境为Python 3.8、搭载2 GHz双核Intel Core i5处理器的Mac,不确定更多CPU是否会改变结果。

测试代码

import concurrent.futures
import multiprocessing as mp
import time
from concurrent.futures.process import ProcessPoolExecutor

import numpy as np

# based on: https://luis-sena.medium.com/sharing-big-numpy-arrays-across-python-processes-abf0dc2a0ab2


def np_sum_global(start, stop):
    return np.sum(data[start:stop])


def benchmark():
    st = time.time()
    ARRAY_SIZE = int(3e8)
    print("array size =", ARRAY_SIZE)
    global data
    data = np.random.random(ARRAY_SIZE)
    print("generated", time.time() - st)
    print("CPU Count =", mp.cpu_count())

    for trial in range(5):
        print("TRIAL =", trial)
        st = time.time()
        s = np.sum(data)
        print("method 1", time.time() - st, s)

        for NUM_WORKERS in range(1, 5):
            st = time.time()
            futures = []
            with ProcessPoolExecutor(max_workers=NUM_WORKERS) as executor:
                for i in range(0, NUM_WORKERS):
                    futures.append(
                        executor.submit(
                            np_sum_global,
                            ARRAY_SIZE * i // NUM_WORKERS,
                            ARRAY_SIZE * (i + 1) // NUM_WORKERS,
                        )
                    )
            futures, _ = concurrent.futures.wait(futures)
            s = sum(future.result() for future in futures)
            print("workers =", NUM_WORKERS, time.time() - st, s)
        print()


if __name__ == "__main__":
    mp.set_start_method("fork")
    benchmark()

测试输出

array size = 300000000
generated 5.1455769538879395
CPU Count = 4
TRIAL = 0
method 1 0.29593801498413086 150004049.39847052
workers = 1 1.8904719352722168 150004049.39847052
workers = 2 1.2082111835479736 150004049.39847034
workers = 3 1.2650330066680908 150004049.39847082
workers = 4 1.233708143234253 150004049.39847046

TRIAL = 1
method 1 0.5861320495605469 150004049.39847052
workers = 1 1.801928997039795 150004049.39847052
workers = 2 1.165492057800293 150004049.39847034
workers = 3 1.2669389247894287 150004049.39847082
workers = 4 1.2941789627075195 150004049.39847043

TRIAL = 2
method 1 0.44912219047546387 150004049.39847052
workers = 1 1.8038971424102783 150004049.39847052
workers = 2 1.1491520404815674 150004049.39847034
workers = 3 1.3324410915374756 150004049.39847082
workers = 4 1.4198641777038574 150004049.39847046

TRIAL = 3
method 1 0.5163640975952148 150004049.39847052
workers = 1 3.248213052749634 150004049.39847052
workers = 2 2.5148861408233643 150004049.39847034
workers = 3 1.0224149227142334 150004049.39847082
workers = 4 1.20924711227417 150004049.39847046

TRIAL = 4
method 1 1.2363107204437256 150004049.39847052
workers = 1 1.8627309799194336 150004049.39847052
workers = 2 1.233341932296753 150004049.39847034
workers = 3 1.3235111236572266 150004049.39847082
workers = 4 1.344843864440918 150004049.39847046

分析与解决方案

为什么multiprocessing比np.sum慢

  • np.sum底层依赖BLAS/LAPACK等高性能计算库(如OpenBLAS、MKL),这些库已经实现了自动多线程优化,在你的双核Mac上会自动利用CPU的多核心能力,无需手动拆分任务。
  • 手动使用multiprocessing会引入额外开销:进程创建、调度、结果通信等成本,完全抵消了并行计算带来的收益,甚至因为CPU资源竞争(子进程的np.sum也会尝试多线程)导致效率更低。

可能的优化方向

  1. 确认numpy的BLAS后端:
    运行np.__config__.show()查看当前numpy使用的计算后端,优先选择MKL或OpenBLAS这类支持多线程的库。如果是单线程后端,可以切换到多线程版本提升np.sum的速度。
  2. 针对超大规模数组(接近内存上限):
    如果数组大小接近系统内存极限,可手动分块读取并求和,避免内存交换(swap)拖慢速度,但只要内存足够,np.sum依然是最优选择。
  3. 避免手动多进程求和:
    对于单纯的数组求和,np.sum的优化已经做到极致,手动多进程完全没必要。只有当计算逻辑复杂、无法被BLAS库优化时,才需要考虑并行化。

更多CPU的影响

如果是核心数更多的机器,np.sum会自动利用更多线程,性能会进一步提升;而手动multiprocessing的开销依然存在,收益依然无法抵消成本,因此还是优先使用np.sum。

内容的提问来源于stack exchange,提问作者b524

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 00:07:02