You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Jupyter Notebook中使用multiprocessing实现函数并行?

在Jupyter Notebook中实现基于进程的并行化方案

问题背景

在Jupyter Notebook中运行CPU密集型任务时,线程池受Python GIL(全局解释器锁)限制无法充分利用多CPU核心,效率低下。尝试使用multiprocessing模块时,因Notebook的运行机制(缺少标准的__main__入口,且.ipynb文件无法被作为Python模块导入)导致报错,且不想创建独立Python模块破坏Notebook的交互式研究流程。

报错复现

最小复现代码

在Notebook单个单元格中运行以下代码:

# 仅用于复现报错,无实际并行逻辑
from multiprocessing import Process

def task():
    return 2

p = Process(target=task)
p.start()
p.join()

不同环境下的报错

  • IPython中运行报错:
Traceback (most recent call last):
  File "<string>", line 1, in <module>
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/multiprocessing/spawn.py", line 116, in spawn_main
    exitcode = _main(fd, parent_sentinel)
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/multiprocessing/spawn.py", line 125, in _main
    prepare(preparation_data)
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/multiprocessing/spawn.py", line 236, in prepare
    _fixup_main_from_path(data['init_main_from_path'])
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/multiprocessing/spawn.py", line 287, in _fixup_main_from_path
    main_content = runpy.run_path(main_path,
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/runpy.py", line 289, in run_path
    return _run_module_code(code, init_globals, run_name,
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/runpy.py", line 96, in _run_module_code
    _run_code(code, mod_globals, init_globals,
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/runpy.py", line 86, in _run_code
    exec(code, run_globals)
  File "/Users/moo/code/ts/trade-executor/notebooks/notebook-multiprocess.ipynb", line 5, in <module>
    "execution_count": null,
NameError: name 'null' is not defined
  • PyCharm/VS Code中运行报错:
Traceback (most recent call last):
  File "<string>", line 1, in <module>
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/multiprocessing/spawn.py", line 116, in spawn_main
    exitcode = _main(fd, parent_sentinel)
  File "/opt/homebrew/Cellar/python@3.10/3.10.13/Frameworks/Python.framework/Versions/3.10/lib/python3.10/multiprocessing/spawn.py", line 126, in _main
    self = reduction.pickle.load(from_parent)
AttributeError: Can't get attribute 'task' on <module '__main__' (built-in)>

解决方案

核心原因是macOS/Linux下Python默认使用spawn启动子进程,该模式需要重新导入主模块,但Notebook的.ipynb文件是JSON格式,无法被正常解析为Python模块。切换到fork启动模式即可解决,fork模式会直接复制父进程的内存空间,子进程无需重新导入模块就能访问父进程中定义的函数。

方法1:直接使用fork上下文创建进程

修改最小复现代码为:

from multiprocessing import Process, get_context

def task():
    return 2

# 切换到fork启动模式
ctx = get_context('fork')
p = ctx.Process(target=task)
p.start()
p.join()

方法2:替换线程池为进程池(适配你的futureproof代码)

将原ThreadPoolExecutor替换为ProcessPoolExecutor,并指定fork上下文:

results = []

def process_background_job(a, b):
   # 处理数据任务逻辑
   pass

import multiprocessing
from futureproof.executors import ProcessPoolExecutor

# 使用fork上下文创建进程池
executor = ProcessPoolExecutor(max_workers=8, mp_context=multiprocessing.get_context('fork'))
with futureproof.TaskManager(executor, error_policy="log") as task_manager:
    
    total_tasks = 0
    for look_back in look_backs:
        for look_forward in look_forwards:
            task_manager.submit(process_background_job, look_back, look_forward)
            total_tasks += 1

    print(f"Processing grid search {total_tasks} background jobs")

    with tqdm(total=total_tasks) as progress_bar:
        for task in task_manager.as_completed():
            if isinstance(task.result, Exception):
                executor.join()
                raise RuntimeError(f"Could not complete task for args {task.args}") from task.result
            
            look_back, look_forward, long_regression, short_regression = task.result
            results.append([
                look_back,
                look_forward,
                long_regression.rsquared,
                short_regression.rsquared
            ])
            progress_bar.update()

注意事项

  • fork模式仅支持Linux/macOS,Windows系统无法使用(Windows仅支持spawn模式)。若需在Windows上使用,可将任务函数定义在独立的.py模块中导入,或使用cloudpickle序列化函数(需额外配置)。
  • 升级Python版本到3.11+不会直接解决该问题,但能提升multiprocessing的稳定性,核心解决方案仍为切换启动模式。

内容的提问来源于stack exchange,提问作者Mikko Ohtamaa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 17:57:57