You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Slurm调度HPC环境下MPI进程派生失败问题求助

在Slurm+HPC环境下通过Python MPI派生Boost.MPI子进程的问题

需求

通过Python脚本(基于mpi4py)在多节点HPC环境中并行运行C子进程(基于Boost.MPI),要求启动多个Python父进程(即mpirun -np y中y>1),每个父进程派生1个C子进程(Spawn的maxprocs=x=1)。目前仅当x=1、y=1时能正常运行,增大x或y均会报错。

相关代码

父进程(Python脚本:mpitest.py)

# parent process : mpitest.py
from mpi4py import MPI

sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=x)

子进程(C++程序:test_mpi.cpp)

// child process: test_mpi.cpp, in which parallelization is implemented using boost::mpi
#include <boost/mpi/environment.hpp>
#include <boost/mpi/communicator.hpp>
#include <boost/mpi.hpp>
#include <iostream>
using namespace boost;

int main(int argc, char* argv[]){
    boost::mpi::environment env(argc, argv);
    boost::mpi::communicator world;
    int commrank;
    MPI_Comm_rank(MPI_COMM_WORLD, &commrank);
    std::cout << commrank << std::endl;
    return 0;
}

初始Slurm脚本

#!/bin/bash
#SBATCH --job-name=mpitest
#SBATCH --partition=day
#SBATCH -N 2
#SBATCH -n 4
#SBATCH -c 6
#SBATCH --mem 5G
#SBATCH -t 01-00:00:00
#SBATCH --output="mpitest.out"
#SBATCH --error="mpitest.error"

#run program
module load Boost/1.74.0-gompi-2020b

#the MPI version is: OpenMPI/4.0.5
mpirun -np y python mpitest.py

错误现象

场景1:x=1、y=2时的报错

[c05n04:13182] pml_ucx.c:178  Error: Failed to receive UCX worker address: Not found (-13)
[c05n04:13182] [[31687,1],1] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493
[c05n11:07942] pml_ucx.c:178  Error: Failed to receive UCX worker address: Not found (-13)
[c05n11:07942] [[31687,2],0] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493
Traceback (most recent call last):
  File "mpitest.py", line 4, in <module>
    sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=1)
  File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn
mpi4py.MPI.Exception: MPI_ERR_OTHER: known error not in list
--------------------------------------------------------------------------
It looks like MPI_INIT failed for some reason; your parallel process is
likely to abort.  There are many reasons that a parallel process can
fail during MPI_INIT; some of which are due to configuration or environment
problems.  This failure appears to be an internal failure; here's some
additional information (which may only be relevant to an Open MPI
developer):

  ompi_dpm_dyn_init() failed
  --> Returned "Error" (-1) instead of "Success" (0)
--------------------------------------------------------------------------
[c05n11:07942] *** An error occurred in MPI_Init
[c05n11:07942] *** reported by process [2076639234,0]
[c05n11:07942] *** on a NULL communicator
[c05n11:07942] *** Unknown error
[c05n11:07942] *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
[c05n11:07942] ***    and potentially your MPI job)

场景2:x=2、y=2时的报错

--------------------------------------------------------------------------
All nodes which are allocated for this job are already filled.
--------------------------------------------------------------------------
Traceback (most recent call last):
  File "mpitest.py", line 4, in <module>
    sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=2)
  File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn
mpi4py.MPI.Exception: MPI_ERR_SPAWN: could not spawn processes
[c18n08:16481] pml_ucx.c:178  Error: Failed to receive UCX worker address: Not found (-13)
[c18n08:16481] [[54742,1],0] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493
[c18n11:01329] pml_ucx.c:178  Error: Failed to receive UCX worker address: Not found (-13)
[c18n11:01329] [[54742,2],0] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493
[c18n11:01332] pml_ucx.c:178  Error: Failed to receive UCX worker address: Not found (-13)
[c18n11:01332] [[54742,2],1] ORTE_ERROR_LOG: Error in file dpm/dpm.c at line 493
Traceback (most recent call last):
  File "mpitest.py", line 4, in <module>
    sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=2)
  File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn
mpi4py.MPI.Exception: MPI_ERR_OTHER: known error not in list
--------------------------------------------------------------------------
It looks like MPI_INIT failed for some reason; your parallel process is
likely to abort.  There are many reasons that a parallel process can
fail during MPI_INIT; some of which are due to configuration or environment
problems.  This failure appears to be an internal failure; here's some
additional information (which may only be relevant to an Open MPI
developer):

  ompi_dpm_dyn_init() failed
  --> Returned "Error" (-1) instead of "Success" (0)
--------------------------------------------------------------------------
[c18n11:01332] *** An error occurred in MPI_Init
[c18n11:01332] *** reported by process [3587571714,1]
[c18n11:01332] *** on a NULL communicator
[c18n11:01332] *** Unknown error
[c18n11:01332] *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
[c18n11:01332] ***    and potentially your MPI job)
[c18n08:16469] 1 more process has sent help message help-mpi-runtime.txt / mpi_init:startup:internal-failure
[c18n08:16469] Set MCA parameter "orte_base_help_aggregate" to 0 to see all help / error messages
[c18n08:16469] 1 more process has sent help message help-mpi-errors.txt / mpi_errors_are_fatal unknown handle

已尝试的修改及新错误

按照建议修改Slurm脚本,禁用UCX相关MCA参数:

#!/bin/bash
#SBATCH --job-name=mpitest
#SBATCH --partition=day
#SBATCH -N 2
#SBATCH -n 4
#SBATCH -c 6
#SBATCH --mem 5G
#SBATCH -t 01-00:00:00
#SBATCH --output="mpitest.out"
#SBATCH --error="mpitest.error"

#run program
module load Boost/1.74.0-gompi-2020b

#the MPI version is: OpenMPI/4.0.5
mpirun --mca pml ^ucx --mca btl ^ucx --mca osc ^ucx -np 2 python mpitest.py

但使用mpirun -np 2且maxprocs=1时,仍报错:

[c18n08:21555] [[49573,1],0] ORTE_ERROR_LOG: Not found in file dpm/dpm.c at line 493
[c18n10:03381] [[49573,3],0] ORTE_ERROR_LOG: Not found in file dpm/dpm.c at line 493
Traceback (most recent call last):
  File "mpitest.py", line 4, in <module>
    sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=1)
  File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn
mpi4py.MPI.Exception: MPI_ERR_INTERN: internal error
[c18n08:21556] [[49573,1],1] ORTE_ERROR_LOG: Not found in file dpm/dpm.c at line 493
[c18n10:03380] [[49573,2],0] ORTE_ERROR_LOG: Not found in file dpm/dpm.c at line 493
Traceback (most recent call last):
  File "mpitest.py", line 4, in <module>
    sub_comm = MPI.COMM_SELF.Spawn('test_mpi', args=[], maxprocs=1)
  File "mpi4py/MPI/Comm.pyx", line 1534, in mpi4py.MPI.Intracomm.Spawn
mpi4py.MPI.Exception: MPI_ERR_INTERN: internal error
--------------------------------------------------------------------------
It looks like MPI_INIT failed for some reason; your parallel process is
likely to abort.  There are many reasons that a parallel process can
fail during MPI_INIT; some of which are due to configuration or environment
problems.  This failure appears to be an internal failure; here's some
additional information (which may only be relevant to an Open MPI
developer):

  ompi_dpm_dyn_init() failed
  --> Returned "Not found" (-13) instead of "Success" (0)
--------------------------------------------------------------------------
[c18n10:03380] *** An error occurred in MPI_Init
[c18n10:03380] *** reported by process [3248816130,0]
[c18n10:03380] *** on a NULL communicator
[c18n10:03380] *** Unknown error
[c18n10:03380] *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
[c18n10:03380] ***    and potentially your MPI job)
[c18n08:21542] 1 more process has sent help message help-mpi-runtime.txt / mpi_init:startup:internal-failure
[c18n08:21542] Set MCA parameter "orte_base_help_aggregate" to 0 to see all help / error messages
[c18n08:21542] 1 more process has sent help message help-mpi-errors.txt / mpi_errors_are_fatal unknown handle

解决方案建议

1. 调整MPI派生的通信子

将MPI.COMM_SELF.Spawn改为MPI.COMM_WORLD.Spawn,避免多个父进程独立派生导致的资源冲突:

# 修改后的父进程代码
from mpi4py import MPI

comm = MPI.COMM_WORLD
# 每个父进程派生1个子进程,总子进程数与父进程数一致
sub_comm = comm.Spawn('test_mpi', args=[], maxprocs=1)

若仅需主进程统一派生所有子进程,可添加root参数:

if comm.rank == 0:
    sub_comm = comm.Spawn('test_mpi', args=[], maxprocs=comm.size, root=0)
else:
    sub_comm = None

2. 调整Slurm与MPI启动参数

  • 确保Slurm分配的任务数(-n)与mpirun -np数值一致,避免资源超配。
  • 添加动态进程管理相关MCA参数,指定使用ORTE管理器:
mpirun --mca ompi_dpm_explicit_launch 1 --mca orte_dyn_launch 1 -np 2 python mpitest.py
  • 若仍有通信层问题,强制使用TCP作为通信后端:
mpirun --mca btl tcp,self --mca pml ob1 -np 2 python mpitest.py

3. 指定子进程绝对路径

派生时传入test_mpi的绝对路径,避免不同节点上找不到可执行文件:

sub_comm = MPI.COMM_WORLD.Spawn('/full/path/to/test_mpi', args=[], maxprocs=1)

4. 验证MPI版本兼容性

确保mpi4py、Boost.MPI与OpenMPI 4.0.5完全兼容,可重新编译mpi4py,绑定当前环境的MPI库。


内容的提问来源于stack exchange,提问作者Formic_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 23:05:30