如何用Python脚本并行提交多组不同输入的test.py任务?
实现test.py的并行执行
首先修正你原代码里的一个小问题:subprocess.run的参数列表中,每个命令行参数应该拆分为单独的元素,比如"-i "+inputs要改成"-i", inputs,否则test.py会把"-i xxx"当成一个完整参数,无法正确解析标记。
下面提供两种常用的并行实现方式:
方法一:使用concurrent.futures.ProcessPoolExecutor
这是Python 3.2+自带的模块,语法简洁,方便控制并发数:
import subprocess from concurrent.futures import ProcessPoolExecutor def run_test_script(input_file): # 拆分参数为独立元素,确保test.py能正确解析 subprocess.run([ 'python', 'test.py', '-i', input_file, '-p', '10', '-o', f"{input_file}_output.csv" ], check=True) if __name__ == '__main__': input_file_list = ["file1", "file2", "file3", "file4", "file5", "file6"] # 你的输入文件列表 # 设置并发数为5,满足你至少5个并行任务的需求 with ProcessPoolExecutor(max_workers=5) as executor: executor.map(run_test_script, input_file_list)
方法二:使用multiprocessing.Pool
这是更传统的多进程池实现,功能同样可靠:
import subprocess from multiprocessing import Pool def run_test_script(input_file): subprocess.run([ 'python', 'test.py', '-i', input_file, '-p', '10', '-o', f"{input_file}_output.csv" ], check=True) if __name__ == '__main__': input_file_list = ["file1", "file2", "file3", "file4", "file5", "file6"] # 初始化进程池,设置5个工作进程 with Pool(processes=5) as pool: pool.map(run_test_script, input_file_list)
注意事项
- 两种方法都是基于多进程,能真正利用多核资源(Python的GIL限制导致多线程更适合IO密集型场景,多进程更适配CPU密集型或外部脚本调用场景)
- 如果你的输入文件数量远大于5,进程池会自动维护5个并行任务,完成一个就补充下一个
check=True会在test.py执行出错时抛出异常,方便排查问题;不需要的话可以去掉- 确保test.py不存在并行执行时的资源竞争问题(你的输出文件是每个输入对应单独的
output.csv,这点已经满足)
内容的提问来源于stack exchange,提问作者user20505088
相关产品推荐
相关产品推荐

