Slurm调度HPC环境下MPI进程派生失败问题求助
在Slurm+HPC环境下通过Python MPI派生Boost.MPI子进程的问题
需求
通过Python脚本(基于mpi4py)在多节点HPC环境中并行运行C子进程(基于Boost.MPI),要求启动多个Python父进程(即mpirun -np y中y>1),每个父进程派生1个C子进程(Spawn的maxprocs=x=1)。目前仅当x=1、y=1时能正常运行,增大x或y均会报错。
相关代码
父进程(Python脚本:mpitest.py)
# parent process : mpitest.py from mpi4py import MPI sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=x)
子进程(C++程序:test_mpi.cpp)
// child process: test_mpi.cpp, in which parallelization is implemented using boost::mpi #include <boost/mpi/environment.hpp> #include <boost/mpi/communicator.hpp> #include <boost/mpi.hpp> #include <iostream> using namespace boost; int main(int argc, char* argv[]){ boost::mpi::environment env(argc, argv); boost::mpi::communicator world; int commrank; MPI_Comm_rank(MPI_COMM_WORLD, &commrank); std::cout << commrank << std::endl; return 0; }
初始Slurm脚本
#!/bin/bash #SBATCH --job-name=mpitest #SBATCH --partition=day #SBATCH -N 2 #SBATCH -n 4 #SBATCH -c 6 #SBATCH --mem 5G #SBATCH -t 01-00:00:00 #SBATCH --output="mpitest.out" #SBATCH --error="mpitest.error" #run program module load Boost/1.74.0-gompi-2020b #the MPI version is: OpenMPI/4.0.5 mpirun -np y python mpitest.py
错误现象
场景1:x=1、y=2时的报错
[c05n04:13182] pml_ucx.c:178 Error: Failed to receive UCX worker address: Not found (-13) [c05n04:13182] [[31687,1],1] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493 [c05n11:07942] pml_ucx.c:178 Error: Failed to receive UCX worker address: Not found (-13) [c05n11:07942] [[31687,2],0] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493 Traceback (most recent call last): File "mpitest.py", line 4, in <module> sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=1) File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn mpi4py.MPI.Exception: MPI_ERR_OTHER: known error not in list -------------------------------------------------------------------------- It looks like MPI_INIT failed for some reason; your parallel process is likely to abort. There are many reasons that a parallel process can fail during MPI_INIT; some of which are due to configuration or environment problems. This failure appears to be an internal failure; here's some additional information (which may only be relevant to an Open MPI developer): ompi_dpm_dyn_init() failed --> Returned "Error" (-1) instead of "Success" (0) -------------------------------------------------------------------------- [c05n11:07942] *** An error occurred in MPI_Init [c05n11:07942] *** reported by process [2076639234,0] [c05n11:07942] *** on a NULL communicator [c05n11:07942] *** Unknown error [c05n11:07942] *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, [c05n11:07942] *** and potentially your MPI job)
场景2:x=2、y=2时的报错
-------------------------------------------------------------------------- All nodes which are allocated for this job are already filled. -------------------------------------------------------------------------- Traceback (most recent call last): File "mpitest.py", line 4, in <module> sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=2) File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn mpi4py.MPI.Exception: MPI_ERR_SPAWN: could not spawn processes [c18n08:16481] pml_ucx.c:178 Error: Failed to receive UCX worker address: Not found (-13) [c18n08:16481] [[54742,1],0] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493 [c18n11:01329] pml_ucx.c:178 Error: Failed to receive UCX worker address: Not found (-13) [c18n11:01329] [[54742,2],0] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493 [c18n11:01332] pml_ucx.c:178 Error: Failed to receive UCX worker address: Not found (-13) [c18n11:01332] [[54742,2],1] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493 Traceback (most recent call last): File "mpitest.py", line 4, in <module> sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=2) File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn mpi4py.MPI.Exception: MPI_ERR_OTHER: known error not in list -------------------------------------------------------------------------- It looks like MPI_INIT failed for some reason; your parallel process is likely to abort. There are many reasons that a parallel process can fail during MPI_INIT; some of which are due to configuration or environment problems. This failure appears to be an internal failure; here's some additional information (which may only be relevant to an Open MPI developer): ompi_dpm_dyn_init() failed --> Returned "Error" (-1) instead of "Success" (0) -------------------------------------------------------------------------- [c18n11:01332] *** An error occurred in MPI_Init [c18n11:01332] *** reported by process [3587571714,1] [c18n11:01332] *** on a NULL communicator [c18n11:01332] *** Unknown error [c18n11:01332] *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, [c18n11:01332] *** and potentially your MPI job) [c18n08:16469] 1 more process has sent help message help-mpi-runtime.txt / mpi_init:startup:internal-failure [c18n08:16469] Set MCA parameter "orte_base_help_aggregate" to 0 to see all help / error messages [c18n08:16469] 1 more process has sent help message help-mpi-errors.txt / mpi_errors_are_fatal unknown handle
已尝试的修改及新错误
按照建议修改Slurm脚本,禁用UCX相关MCA参数:
#!/bin/bash #SBATCH --job-name=mpitest #SBATCH --partition=day #SBATCH -N 2 #SBATCH -n 4 #SBATCH -c 6 #SBATCH --mem 5G #SBATCH -t 01-00:00:00 #SBATCH --output="mpitest.out" #SBATCH --error="mpitest.error" #run program module load Boost/1.74.0-gompi-2020b #the MPI version is: OpenMPI/4.0.5 mpirun --mca pml ^ucx --mca btl ^ucx --mca osc ^ucx -np 2 python mpitest.py
但使用mpirun -np 2且maxprocs=1时,仍报错:
[c18n08:21555] [[49573,1],0] ORTE_ERROR_LOG: Not found in file dpm/dpm.c at line 493 [c18n10:03381] [[49573,3],0] ORTE_ERROR_LOG: Not found in file dpm/dpm.c at line 493 Traceback (most recent call last): File "mpitest.py", line 4, in <module> sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=1) File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn mpi4py.MPI.Exception: MPI_ERR_INTERN: internal error [c18n08:21556] [[49573,1],1] ORTE_ERROR_LOG: Not found in file dpm/dpm.c at line 493 [c18n10:03380] [[49573,2],0] ORTE_ERROR_LOG: Not found in file dpm/dpm.c at line 493 Traceback (most recent call last): File "mpitest.py", line 4, in <module> sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=1) File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn mpi4py.MPI.Exception: MPI_ERR_INTERN: internal error -------------------------------------------------------------------------- It looks like MPI_INIT failed for some reason; your parallel process is likely to abort. There are many reasons that a parallel process can fail during MPI_INIT; some of which are due to configuration or environment problems. This failure appears to be an internal failure; here's some additional information (which may only be relevant to an Open MPI developer): ompi_dpm_dyn_init() failed --> Returned "Not found" (-13) instead of "Success" (0) -------------------------------------------------------------------------- [c18n10:03380] *** An error occurred in MPI_Init [c18n10:03380] *** reported by process [3248816130,0] [c18n10:03380] *** on a NULL communicator [c18n10:03380] *** Unknown error [c18n10:03380] *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, [c18n10:03380] *** and potentially your MPI job) [c18n08:21542] 1 more process has sent help message help-mpi-runtime.txt / mpi_init:startup:internal-failure [c18n08:21542] Set MCA parameter "orte_base_help_aggregate" to 0 to see all help / error messages [c18n08:21542] 1 more process has sent help message help-mpi-errors.txt / mpi_errors_are_fatal unknown handle
解决方案建议
1. 调整MPI派生的通信子
将MPI.COMM_SELF.Spawn改为MPI.COMM_WORLD.Spawn,避免多个父进程独立派生导致的资源冲突:
# 修改后的父进程代码 from mpi4py import MPI comm = MPI.COMM_WORLD # 每个父进程派生1个子进程,总子进程数与父进程数一致 sub_comm = comm.Spawn('test_mpi', args=[], maxprocs=1)
若仅需主进程统一派生所有子进程,可添加root参数:
if comm.rank == 0: sub_comm = comm.Spawn('test_mpi', args=[], maxprocs=comm.size, root=0) else: sub_comm = None
2. 调整Slurm与MPI启动参数
- 确保Slurm分配的任务数(
-n)与mpirun -np数值一致,避免资源超配。 - 添加动态进程管理相关MCA参数,指定使用ORTE管理器:
mpirun --mca ompi_dpm_explicit_launch 1 --mca orte_dyn_launch 1 -np 2 python mpitest.py
- 若仍有通信层问题,强制使用TCP作为通信后端:
mpirun --mca btl tcp,self --mca pml ob1 -np 2 python mpitest.py
3. 指定子进程绝对路径
派生时传入test_mpi的绝对路径,避免不同节点上找不到可执行文件:
sub_comm = MPI.COMM_WORLD.Spawn('/full/path/to/test_mpi', args=[], maxprocs=1)
4. 验证MPI版本兼容性
确保mpi4py、Boost.MPI与OpenMPI 4.0.5完全兼容,可重新编译mpi4py,绑定当前环境的MPI库。
内容的提问来源于stack exchange,提问作者Formic_
相关产品推荐
相关产品推荐

