You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ProcessPoolExecutor并行运行CPU密集型脚本时的两个问题:无法并行及子进程无打印输出

ProcessPoolExecutor并行运行CPU密集型脚本时的两个问题:无法并行及子进程无打印输出

一、为什么最初的代码无法实现并行?

你最初的代码里有两个关键错误,直接导致了脚本串行执行:

  1. 提前在主进程执行了函数:你用exec(f'{script}.run()')这行代码时,会直接在主进程里立刻调用对应脚本的run()函数,完全没交给ProcessPoolExecutor去调度。所以脚本1先跑完3秒,脚本2才开始跑4秒,总时间自然是7秒。
  2. executor.map的用法错误:map方法需要接收一个可调用对象+参数序列,而不是直接传入函数执行后的结果,你的写法完全偏离了它的正确使用方式。

后来你改用executor.submit(script)的方式就对了——直接把函数对象提交给进程池,由子进程去执行,这样才能真正实现并行。

二、为什么ProcessPoolExecutor下子进程的print没有输出?

这个问题在Spyder环境里特别常见,核心原因是ProcessPoolExecutor创建的子进程,其标准输出(stdout)被系统缓冲了,或者Spyder的控制台没有正确捕获子进程的输出流;而ThreadPoolExecutor用的是主线程的输出流,所以能正常打印内容。

你可以试试这几个解决办法:

1. 强制刷新子进程的输出流

修改两个测试脚本里的print语句,加上flush=True参数,强制立即输出内容,不让系统缓冲:

# test_script_1.py里的print修改后
print('starting script 1', flush=True)
print(f'Running script 1: Sleeping {t} second(s)', flush=True)

# test_script_2.py同理
print('starting script 2', flush=True)
print(f'Sleeping {t} second(s)', flush=True)

2. 调整Spyder的运行设置

在Spyder里可以尝试两种设置:

  • 打开Tools > Preferences > IPython console > Graphics,把后端改成Qt5(避开Automatic/Inline)
  • 打开Tools > Preferences > Run,勾选Execute in an external system terminal,让脚本在外部终端运行,子进程的输出就能正常显示了

3. 手动用队列捕获子进程输出(进阶方案)

如果上面的方法都不行,可以用multiprocessing.Queue让子进程把输出内容传给主进程,再由主进程打印:

import time
from concurrent.futures import ProcessPoolExecutor
import multiprocessing
import test_script_1, test_script_2

def wrapper(func, queue):
    # 把子进程的stdout重定向到队列
    import sys
    class QueueWriter:
        def __init__(self, q):
            self.q = q
        def write(self, msg):
            self.q.put(msg)
        def flush(self):
            pass
    sys.stdout = QueueWriter(queue)
    func()

if __name__ == '__main__':
    start = time.perf_counter()
    print("Starting with scripts execution")
    queue = multiprocessing.Queue()
    
    script_list = [test_script_1.run, test_script_2.run]
    with ProcessPoolExecutor() as executor:
        futures = [executor.submit(wrapper, script, queue) for script in script_list]
        
        # 循环读取队列并打印子进程输出
        while any(future.running() for future in futures) or not queue.empty():
            while not queue.empty():
                print(queue.get(), end='')
    
    print('all scripts done')
    finish = time.perf_counter()
    print(f'Finished in {(finish - start):.2f} seconds')

总结

  • 要实现并行,必须确保把函数对象提交给进程池,绝对不能在主进程里提前执行函数
  • ProcessPoolExecutor的子进程输出问题,大多和输出缓冲或环境的控制台捕获机制有关,优先尝试刷新输出流或调整Spyder运行设置

备注:内容来源于stack exchange,提问作者Guido De Haan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 08:08:11