Python concurrent.futures多进程仅用两核心,如何全核心利用?
问题分析与解决
你的代码里有个关键错误:调用executor.submit时直接执行了process_single函数,而不是把函数和参数传给进程池异步执行。这种写法会导致所有预处理任务都在主线程串行运行,完全没用到多进程,自然CPU核心利用率上不去。
修正步骤
1. 正确使用submit方法
submit需要接收函数对象作为第一个参数,后面跟函数的参数,而不是直接调用函数。修正后的代码如下:
import concurrent.futures import os import numpy as np def preprocessing(tar_ratio, img_paths, label_paths, save_dir="output", resampling_mode=None): # 显式指定max_workers为CPU核心数,默认也是os.cpu_count(),显式写更清晰 with concurrent.futures.ProcessPoolExecutor(max_workers=os.cpu_count()) as executor: futures = [] for img_path, label_path in zip(img_paths, label_paths): src_ratio = get_ratio(label_path) if not np.isnan(src_ratio): # 只传函数名,参数依次放在后面 future = executor.submit( process_single, src_ratio, tar_ratio, img_path, label_path, save_dir=save_dir, resampling_mode=resampling_mode ) futures.append(future) # 可选:等待所有任务完成再退出 concurrent.futures.wait(futures)
2. 更简洁的map用法(推荐)
如果是批量处理可迭代的输入,用starmap比submit更简洁,它会自动分配任务到进程池:
import concurrent.futures import numpy as np def preprocessing(tar_ratio, img_paths, label_paths, save_dir="output", resampling_mode=None): # 先过滤掉无效的ratio数据 valid_tasks = [ (get_ratio(label_path), tar_ratio, img_path, label_path, save_dir, resampling_mode) for img_path, label_path in zip(img_paths, label_paths) if not np.isnan(get_ratio(label_path)) ] with concurrent.futures.ProcessPoolExecutor() as executor: # starmap会把元组里的元素逐个传给process_single executor.starmap(process_single, valid_tasks)
核心注意事项
- ProcessPoolExecutor默认行为:不指定
max_workers时,会自动使用os.cpu_count()的值(即CPU总核心数),能最大化利用CPU资源。 - 禁止提前执行函数:这是新手常犯的错误,必须确保传给
submit的是函数对象,而非函数执行后的返回值。 - CPU密集型任务选型逻辑:用
ProcessPoolExecutor是正确的,Python的GIL会限制多线程在CPU密集型任务中的性能,多进程能绕过GIL限制。
内容的提问来源于stack exchange,提问作者Zheng
相关产品推荐
相关产品推荐

