如何用concurrent.futures将FileList元素分配给CPU并行运行Python脚本?
用concurrent.futures并行执行多个Python脚本
没问题,我来帮你把串行执行的脚本改成并行版本!concurrent.futures.ProcessPoolExecutor正好适合这个场景——它会自动把任务分配到不同的CPU核心,完全不用你手动操心核心分配的事儿。
先给你理清楚思路,再上代码:
核心思路
- 先封装一个单独执行单个Python脚本的函数,把执行逻辑收束起来(用
subprocess启动独立Python进程执行脚本,和你原来的串行逻辑保持一致); - 用
ProcessPoolExecutor创建进程池,把FileList里的每个脚本路径批量丢给进程池,它会自动调度到空闲的CPU核心去执行。
代码示例
先看你的串行版本(参考)
假设你原来的串行代码大概是这样:
import subprocess import os # 假设这是你的FileList,包含所有要执行的Python脚本路径 file_list = ["./script1.py", "./script2.py", "./script3.py", "./script4.py"] def run_single_script(file_path): # 转成绝对路径,避免工作目录带来的路径问题 abs_path = os.path.abspath(file_path) print(f"开始执行: {abs_path}") try: # 调用Python解释器执行脚本 result = subprocess.run( ["python", abs_path], check=True, capture_output=True, text=True ) print(f"{abs_path} 执行成功,输出:\n{result.stdout}") except subprocess.CalledProcessError as e: print(f"{abs_path} 执行失败,错误:\n{e.stderr}") # 串行循环执行 for script in file_list: run_single_script(script)
改成并行版本只需要几步
把串行的循环替换成ProcessPoolExecutor批量提交任务:
import subprocess import os from concurrent.futures import ProcessPoolExecutor file_list = ["./script1.py", "./script2.py", "./script3.py", "./script4.py"] def run_single_script(file_path): abs_path = os.path.abspath(file_path) print(f"开始执行: {abs_path}") try: result = subprocess.run( ["python", abs_path], check=True, capture_output=True, text=True ) return (abs_path, "success", result.stdout) except subprocess.CalledProcessError as e: return (abs_path, "failed", e.stderr) if __name__ == "__main__": # 进程池大小默认是CPU核心数,也可以手动指定比如max_workers=4 with ProcessPoolExecutor() as executor: # map方法自动把任务分配到不同进程,不用手动管核心 results = executor.map(run_single_script, file_list) # 统一处理所有任务的执行结果 for res in results: script_path, status, output = res if status == "success": print(f"\n{script_path} 执行成功:") print(output) else: print(f"\n{script_path} 执行失败:") print(output)
关键细节说明
- 为什么选ProcessPoolExecutor?:执行Python脚本属于CPU密集型任务,线程池会受GIL(全局解释器锁)限制,而进程池是真正的多进程并行,能充分利用多个CPU核心。
- 自动分配核心:你不用手动给每个脚本指定核心,
ProcessPoolExecutor底层会维护进程池,自动把任务调度到空闲进程(对应CPU核心)执行。 - 异常隔离:把执行结果封装成返回值,就算某个脚本执行失败,也不会影响其他脚本的运行。
- Windows必加
if __name__ == "__main__":Windows系统下多进程启动时需要重新导入模块,加这个判断能避免无限递归创建进程。
如果你的脚本需要传递参数,只需要修改函数和任务列表:
def run_single_script(args): file_path, arg1, arg2 = args abs_path = os.path.abspath(file_path) # 执行时带上参数 result = subprocess.run( ["python", abs_path, arg1, arg2], check=True, capture_output=True, text=True ) # 后续逻辑不变... # 把任务改成带参数的元组列表 task_list = [("./script1.py", "param1", "param2"), ("./script2.py", "param3", "param4")] with ProcessPoolExecutor() as executor: results = executor.map(run_single_script, task_list)
内容的提问来源于stack exchange,提问作者Angel_M
相关产品推荐
相关产品推荐

