You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Open MPI发送超特定大小数组时出现段错误

Open MPI多节点大消息通信段错误排查与解决

问题背景

运行于InfiniBand+Intel架构集群,采用Open MPI、PMIx及SLURM调度的程序,当单节点输入矩阵规模超过38×38(对应示例中x>1500)时,跨节点的发送/接收、集合调用触发段错误;单节点或使用Intel MPI时无异常,仅多节点搭配Open MPI场景下出现问题。程序虽能执行完毕,但错误会阻塞后续任务启动。

复现代码

int main(int argc, char** argv) {
    int proc_size, p_rank;

    MPI_Init(&argc, &argv);
    MPI_Comm_size(MPI_COMM_WORLD, &proc_size);
    MPI_Comm_rank(MPI_COMM_WORLD, &p_rank); 

    int x = 1600;

    MPI_Status status;
    double* A = calloc(x, sizeof(double));

    if (p_rank == 0)
        for (int i = 1; i < proc_size; ++i) {
            MPI_Recv(A, x, MPI_DOUBLE, i, 0, MPI_COMM_WORLD, &status);
        }
    else
        MPI_Send(A, x, MPI_DOUBLE, 0, 0, MPI_COMM_WORLD);

    MPI_Finalize();

    return 0;
}

执行方式与错误信息

使用SBatch作业执行:

srun --mpi=pmix -n 36 my_program

或直接用mpirun执行,均会触发段错误,错误日志示例:

[node30:14535] *** Process received signal ***
[node42:144621] *** Process received signal ***
[node42:144621] Signal: Segmentation fault (11)
[node42:144621] Signal code: Address not mapped (1)
[node42:144621] Failing at address: 0x7fcc6fdcc210
[node30:14535] Signal: Segmentation fault (11)
[node30:14535] Signal code: Address not mapped (1)
[node30:14535] Failing at address: 0x7fe9fe17d210
[node19:91882] *** Process received signal ***
[node19:91882] Signal: Segmentation fault (11)
[node19:91882] Signal code: Address not mapped (1)
[node19:91882] Failing at address: 0x7fb02739d210
srun: error: node30: task 4: Segmentation fault
srun: error: node42: task 7: Segmentation fault

排查与解决方案

1. 检查内存锁定限制

InfiniBand的RDMA传输需要锁定内存以避免缓冲区被换出,若系统内存锁定限制不足会导致段错误:

  • 在SBatch脚本中添加内存锁定配置:
    ulimit -l unlimited
    
  • 长期修改可编辑/etc/security/limits.conf,添加:
    * soft memlock unlimited
    * hard memlock unlimited
    

2. 调整Open MPI运行参数

强制启用InfiniBand传输并配置内存锁定:

srun --mpi=pmix -n 36 --mca btl_openib_allow_ib 1 --mca btl_openib_mem_lock 1 my_program

或添加--mca mpi_leave_pinned 1让MPI保持内存固定:

srun --mpi=pmix -n 36 --mca mpi_leave_pinned 1 my_program

3. 使用MPI专用内存分配

替换calloc为MPI_Alloc_mem,该函数会自动将内存注册到RDMA设备,避免手动注册问题:

int main(int argc, char** argv) {
    int proc_size, p_rank;

    MPI_Init(&argc, &argv);
    MPI_Comm_size(MPI_COMM_WORLD, &proc_size);
    MPI_Comm_rank(MPI_COMM_WORLD, &p_rank); 

    int x = 1600;

    MPI_Status status;
    double* A;
    // 使用MPI分配内存并初始化
    MPI_Alloc_mem(x * sizeof(double), MPI_INFO_NULL, &A);
    memset(A, 0, x * sizeof(double));

    if (p_rank == 0)
        for (int i = 1; i < proc_size; ++i) {
            MPI_Recv(A, x, MPI_DOUBLE, i, 0, MPI_COMM_WORLD, &status);
        }
    else
        MPI_Send(A, x, MPI_DOUBLE, 0, 0, MPI_COMM_WORLD);

    MPI_Free_mem(A); // 释放MPI分配的内存
    MPI_Finalize();

    return 0;
}

4. 升级Open MPI版本

旧版本Open MPI可能存在InfiniBand与PMIx交互的bug,升级到4.x系列最新稳定版可修复部分已知问题。

5. 检查SLURM资源分配

确保SBatch脚本正确指定节点与任务分布,例如36任务分2节点:

#SBATCH --nodes=2
#SBATCH --ntasks-per-node=18

避免任务跨节点分配时的资源冲突。

内容的提问来源于stack exchange,提问作者Another Shrubbery

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 13:48:13