You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MPI_Bcast报'failed to attach to a bootstrap queue'错误排查

MPI程序处理1000万+元素时崩溃的问题排查与修复

我编写了一个运行在4节点MPI环境下的C程序,流程为接收整数N,通过MPI_Bcast广播至所有节点,每个节点动态创建大小为N的数组。当N在64到100万之间时程序运行正常,但输入1000万及以上元素时,MPI会崩溃,偶尔出现如下错误:

Fatal error in MPI_Bcast: Other MPI error, error stack:
MPI_Bcast(buf=0x000000000067FD74, count=1, MPI_INT, root=0, MPI_COMM_WORLD) failed
failed to attach to a bootstrap queue - 6664:280

1000万在int的取值范围内,不确定错误原因,相关代码如下:

#include <stdio.h>
#include <stdlib.h>
#include "mpi.h"
#include <time.h>

int main(int argc, char *argv[]){

    int process_Rank, size_Of_Cluster;
    
    int number_of_elements;

    MPI_Init(&argc, &argv);
    MPI_Comm_size(MPI_COMM_WORLD, &size_Of_Cluster);
    MPI_Comm_rank(MPI_COMM_WORLD, &process_Rank);

    if(process_Rank == 0){
        printf("Enter the number of elements:\n");
        fflush(stdout);
        scanf("%d", &number_of_elements);   
    }

    MPI_Bcast(&number_of_elements,1,MPI_INT, 0, MPI_COMM_WORLD);

    int *outputs = (int*)malloc(number_of_elements * sizeof(int));

    unsigned long long chunk_size = number_of_elements/ size_Of_Cluster;

    int my_input[chunk_size], my_output[chunk_size];

    for(int i = 0; i < number_of_elements; i++){
        outputs[i] = i+1;
    }

    MPI_Barrier(MPI_COMM_WORLD);

    clock_t begin = clock();

    MPI_Scatter(outputs, chunk_size, MPI_INT, &my_input, chunk_size, MPI_INT, 0, MPI_COMM_WORLD);

    for(int i = 0; i <= chunk_size; i++){
        my_output[i] = my_input[i];
    }
    
    MPI_Gather(&my_output, chunk_size, MPI_INT, outputs, chunk_size, MPI_INT, 0, MPI_COMM_WORLD);

    int iterate_terms[5] = {2,4,8,4,2};
    int starting_terms[5] = {1,3,7,3,1};
    int subtract_terms[5] = {1,2,4,0,0};
    int adding_terms[5] = {0,0,0,2,1};

    for(int j = 0; j < 5; j++){

        MPI_Scatter(outputs, chunk_size, MPI_INT, &my_input, chunk_size, MPI_INT, 0, MPI_COMM_WORLD);

        for(int i = starting_terms[j]; i <= chunk_size; i+= iterate_terms[j]){
            my_output[i+adding_terms[j]] += my_input[i-subtract_terms[j]];
        }
         
        MPI_Gather(&my_output, chunk_size, MPI_INT, outputs, chunk_size, MPI_INT, 0, MPI_COMM_WORLD);
    }

    MPI_Barrier(MPI_COMM_WORLD);

    if(process_Rank == 0){
        for(int i = chunk_size-1; i < number_of_elements; i+=chunk_size){
            outputs[i+1] += outputs[i];
            outputs[i+2] += outputs[i];
            outputs[i+3] += outputs[i];
        }

        clock_t end = clock();

        double time_spent = (double)(end-begin) / CLOCKS_PER_SEC;

        for(int i = 0; i < number_of_elements; i++){
            printf("%d \n", outputs[i]);
            fflush(stdout);
        }

        printf("took %f", time_spent);
        fflush(stdout);
    } else {
        clock_t end = clock();
    }

    MPI_Finalize();
    return 0;
}

核心问题:栈溢出引发的进程崩溃

错误提示指向MPI_Bcast,但实际根源是栈上分配的变长数组(VLA)超出了进程默认栈大小限制。

代码中这两行是关键问题点:

unsigned long long chunk_size = number_of_elements/ size_Of_Cluster;
int my_input[chunk_size], my_output[chunk_size];

当N=1000万时,4节点下chunk_size为250万,每个int占4字节,单个数组就需要10MB内存,两个数组合计20MB。而多数操作系统的默认栈大小仅为8-10MB,栈上分配这么大的数组会直接触发栈溢出,导致进程异常终止,进而引发MPI通信错误(进程崩溃后无法完成广播操作)。

修复方案

将栈上的变长数组改为堆上动态分配,避免栈溢出:

unsigned long long chunk_size = number_of_elements/ size_Of_Cluster;
// 替换栈上数组为堆分配
int *my_input = (int*)malloc(chunk_size * sizeof(int));
int *my_output = (int*)malloc(chunk_size * sizeof(int));

// 程序结束前释放内存
if (my_input != NULL) free(my_input);
if (my_output != NULL) free(my_output);

其他潜在问题修复

  1. 检查内存分配是否成功:动态分配后必须判断是否分配成功,避免空指针访问:
int *outputs = (int*)malloc(number_of_elements * sizeof(int));
if (outputs == NULL) {
    fprintf(stderr, "Process %d: 分配outputs数组失败\n", process_Rank);
    MPI_Abort(MPI_COMM_WORLD, 1);
}
  1. 修复循环越界问题:代码中多处循环条件错误,比如:
for(int i = 0; i <= chunk_size; i++){
    my_output[i] = my_input[i];
}

数组下标范围是0到chunk_size-1,应改为i < chunk_size,否则会访问越界内存。同理,后续循环中也要确保i+adding_terms[j]不超过chunk_size-1。

  1. 处理N无法被节点数整除的情况:当前代码假设N能被4整除,若不能,最后一个节点的chunk_size会更小,需改用MPI_Scatterv和MPI_Gatherv处理不等长数据分发,避免通信出错。

总结

直接导致崩溃的原因是栈上大数组引发的栈溢出,改成堆分配即可解决核心问题。同时修复循环越界、内存分配检查等问题,能大幅提升程序健壮性。

内容的提问来源于stack exchange,提问作者Matthew Haywood

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 18:05:33