You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Multiprocessing/Multithreading并行化Python for循环

多进程并行优化:解决双参数函数循环耗时问题

核心问题拆解

你之前遇到的问题本质是对多进程的内存模型和参数传递逻辑不熟悉:

  • 多进程不共享全局内存:每个子进程会复制主进程的内存空间,修改子进程里的output_list不会同步到主进程,这是列表为空的根本原因
  • 参数构造错误:之前的input_args = [(list1, list2) for x in list2]是把整个list2传给每个任务,而非单个元素,导致任务计算量翻倍
  • 函数依赖全局变量:无论是主文件还是单独文件的函数,依赖全局output_list都会因为进程隔离导致变量未定义或修改无效

正确实现方案

方案1:使用multiprocessing.Pool.starmap

适合需要传递多参数的场景,自动拆解元组作为函数参数,同时用if __name__ == '__main__':避免多进程启动时的递归导入错误(解决之前的AttributeError)。

from multiprocessing import Pool

# 纯函数:只接收参数、计算、返回结果,不依赖全局变量
def my_function(list1, item_from_list2):
    # 替换为你的实际计算逻辑
    # do_stuff
    # do_other_stuff_to_get_statistical_score_between_list1_and_item_from_list2
    return score

if __name__ == '__main__':
    list1 = [a,b,c,d,e,f,g]  # 你的实际数据
    list2 = [1,2,3,4,5,6,7,8,9]
    
    # 构造正确的参数对:每个list2元素和list1一一配对
    input_args = [(list1, x) for x in list2]
    
    # 启动进程池,指定CPU核心数(比如4个)
    with Pool(processes=4) as pool:
        # starmap自动拆解每个元组,传入my_function
        output_list = pool.starmap(my_function, input_args)
    
    # output_list是按list2顺序排列的所有计算结果
    print(output_list)

方案2:使用concurrent.futures.ProcessPoolExecutor(更简洁API)

用functools.partial固定第一个参数list1,让每个任务只需要传递list2的单个元素,代码更简洁。

from concurrent.futures import ProcessPoolExecutor
from functools import partial

def my_function(list1, item_from_list2):
    # 替换为你的实际计算逻辑
    return score

if __name__ == '__main__':
    list1 = [a,b,c,d,e,f,g]
    list2 = [1,2,3,4,5,6,7,8,9]
    
    # 固定list1参数,生成只接受单个参数的新函数
    func_with_list1 = partial(my_function, list1)
    
    with ProcessPoolExecutor(max_workers=4) as executor:
        # map遍历list2,将每个元素传入处理后的函数
        output_list = list(executor.map(func_with_list1, list2))
    
    print(output_list)

为什么之前的并行代码更慢?

  1. 参数错误:把整个list2传给每个任务,导致每个任务重复计算全量数据,计算量远超串行
  2. 全局变量开销:多进程中依赖全局变量会触发额外的进程间通信开销,甚至锁竞争
  3. 任务粒度问题:如果单个任务计算时间极短,进程创建/销毁的开销会超过并行收益,此时CPU密集型任务建议合并小任务,IO密集型任务改用多线程(ThreadPoolExecutor)

函数单独存放的正确姿势

如果必须把my_function放在单独文件(比如sams_functions.py),确保函数只通过参数接收数据,不依赖全局变量:

sams_functions.py:

def my_function(list1, item_from_list2):
    # 替换为你的实际计算逻辑
    return score

主文件:

from concurrent.futures import ProcessPoolExecutor
from functools import partial
from sams_functions import my_function

if __name__ == '__main__':
    list1 = [a,b,c,d,e,f,g]
    list2 = [1,2,3,4,5,6,7,8,9]
    
    func_partial = partial(my_function, list1)
    with ProcessPoolExecutor(max_workers=4) as executor:
        output_list = list(executor.map(func_partial, list2))

内容的提问来源于stack exchange,提问作者Rainman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 08:23:22