Jupyter Notebook使用multiprocessing运行代码无限循环问题求助
问题原因
你遇到的卡死问题是multiprocessing的启动逻辑和Jupyter交互环境特性冲突导致的:
- Windows、macOS平台下Python的multiprocessing默认使用spawn方式启动子进程,子进程启动时会重新导入当前运行的全量代码
- 你在全局作用域直接创建
mp.Pool执行任务,没有添加if __name__ == '__main__'保护逻辑,子进程导入代码时会重复创建进程池、提交任务,最终陷入无限循环 - Jupyter交互环境的代码没有明确的模块入口,默认所有代码都在
__main__作用域运行,进一步放大了这个问题
解决方案
方法1:添加进程保护逻辑
将进程池创建、任务提交的代码放到if __name__ == '__main__'代码块中,同时手动回收进程资源,修改后代码如下:
import multiprocessing as mp import numpy as np def square(x): return np.square(x) if __name__ == '__main__': x = np.arange(64) pool = mp.Pool(4) squared = pool.map(square, [x[16*i:16*i+16] for i in range(4)]) pool.close() pool.join() # 可按需打印结果验证 print(squared)
方法2:使用Jupyter适配性更好的进程池实现
concurrent.futures.ProcessPoolExecutor对Jupyter环境的兼容表现更好,且支持上下文自动回收资源,无需手动处理进程关闭逻辑,代码示例如下:
from concurrent.futures import ProcessPoolExecutor import numpy as np def square(x): return np.square(x) x = np.arange(64) with ProcessPoolExecutor(4) as executor: squared = list(executor.map(square, [x[16*i:16*i+16] for i in range(4)]))
额外注意事项
如果上述方法依然运行异常,可以尝试将工作函数square单独写到独立的.py文件中再导入使用,完全规避Jupyter作用域的影响;不要在子进程运行的函数中调用Jupyter专属的魔法命令、交互变量,会导致子进程执行失败。
内容的提问来源于stack exchange,提问作者vageesh
相关产品推荐
相关产品推荐

