You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python multiprocessing并行求和性能不升反降的问题咨询

Python multiprocessing并行求和性能不升反降的问题咨询

我的实验情况与困惑

为了搞懂并行化,我基于multiprocessing模块写了个求和的玩具脚本,结果却发现多进程反而让性能变得更差,完全和我预期的相反。

我生成了一个包含10000个元素的列表,尝试用多进程并行求和:每个进程负责计算自己分到的子列表的和,最后汇总结果。但测试结果扎心了:

  • 1个进程:耗时0.78秒
  • 2个进程:耗时1.29秒
  • 3个进程:耗时1.93秒

我换了两种不同的硬件测试,也试了不同数量的进程,结果都是多进程耗时更长。我本来以为固定列表大小下,每个进程只处理一小部分数据,并行执行应该能减少总时间,现在完全搞不懂问题出在哪了?是不是因为多个进程都要写入同一个result变量,互相通信导致性能下降?

我的测试代码

import multiprocessing
import time
import math
import numpy as np

def sum_list(mylist,result,index_min,index_max):
    for i in range(index_min,index_max):
        result.value += mylist[i]

if __name__ == "__main__":
    f = lambda x: np.sin(x)*np.cos(x)
    X = np.linspace(0,1,10000)
    mylist=f(X)  # 创建要求和的列表
    size = len(mylist)
    
    result = multiprocessing.Value('d')
    result.value = 0
    
    processes = []
    n = 10  # 我测试过不同的n值,比如1、2、3等
    
    # 创建进程
    start = time.perf_counter()
    for p in range(n):
        index_min = int(p*size/n)
        index_max = int((p+1)*size/n)
        print(index_min,index_max)
        processes.append( multiprocessing.Process(target = sum_list,args = (mylist,result,index_min,index_max)) )
    
    # 启动进程
    for process in processes:
        process.start()
    
    # 等待进程结束
    for process in processes:
        process.join()
    
    end = time.perf_counter()
    print('time ellapsed :', end-start, 'seconds. ')

问题分析与改进建议

你猜的没错,核心问题确实和共享变量的竞争有关,另外还有几个并行化新手容易踩的坑:

  1. 共享变量的锁开销拉满
    你用的multiprocessing.Value是个进程间同步的共享对象,每次对result.value做+=操作时,底层都会触发锁的获取和释放——多个进程挤着抢这一把锁,大部分时间都在等锁,根本没真正并行计算。就像一群人抢一个计算器用,反而比一个人慢慢算还慢。

  2. 进程本身的开销盖过了并行收益
    每个子进程启动时,会拷贝父进程的内存空间(Linux下是写时拷贝,但这里的mylist是numpy数组,子进程访问时可能触发实际拷贝),加上进程启动、调度的成本,对于“求和10000个元素”这种计算量极小的任务来说,这些开销完全把并行的优势给吞了。

  3. 任务粒度太小,并行优势无从谈起
    10000个元素分给10个进程,每个进程只处理1000个元素,计算时间可能连1毫秒都不到,而进程启动、锁竞争的时间都是毫秒级的,并行的优势根本发挥不出来。


具体的改进方案

方案1:去掉共享变量,让进程独立计算后汇总

让每个进程自己算负责的子列表的和,最后在主进程把所有子和加起来,完全避免进程间的锁竞争:

import multiprocessing
import time
import numpy as np

def sum_sublist(mylist, index_min, index_max):
    sub_sum = 0.0
    for i in range(index_min, index_max):
        sub_sum += mylist[i]
    return sub_sum

if __name__ == "__main__":
    f = lambda x: np.sin(x)*np.cos(x)
    X = np.linspace(0,1,10000)
    mylist = f(X)
    size = len(mylist)
    
    n = 10
    start = time.perf_counter()
    
    # 用Pool管理进程更省心
    with multiprocessing.Pool(processes=n) as pool:
        tasks = []
        for p in range(n):
            idx_min = int(p * size / n)
            idx_max = int((p+1) * size / n)
            tasks.append((mylist, idx_min, idx_max))
        # 并行执行,拿到所有子和
        sub_results = pool.starmap(sum_sublist, tasks)
    
    total_sum = sum(sub_results)
    end = time.perf_counter()
    
    print('总求和结果:', total_sum)
    print('耗时:', end - start, '秒')

方案2:直接用numpy内置求和,比手动循环快N倍

其实对于numpy数组来说,mylist.sum()是底层C实现的,比Python循环+并行快得多,甚至不需要折腾并行:

start = time.perf_counter()
print('Numpy求和结果:', mylist.sum())
print('耗时:', time.perf_counter() - start, '秒')

方案3:增大计算量,让并行优势显现

如果一定要测试并行的效果,把列表元素数量改成1000万甚至更多,或者给每个元素加更复杂的计算(比如多嵌套几次三角函数运算),让计算时间远大于进程启动和通信的开销,这时并行的优势才会体现出来。


备注:内容来源于stack exchange,提问作者Mi-Shell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 08:58:00