MPI_Bcast报'failed to attach to a bootstrap queue'错误排查
我编写了一个运行在4节点MPI环境下的C程序,流程为接收整数N,通过MPI_Bcast广播至所有节点,每个节点动态创建大小为N的数组。当N在64到100万之间时程序运行正常,但输入1000万及以上元素时,MPI会崩溃,偶尔出现如下错误:
Fatal error in MPI_Bcast: Other MPI error, error stack:
MPI_Bcast(buf=0x000000000067FD74, count=1, MPI_INT, root=0, MPI_COMM_WORLD) failed
failed to attach to a bootstrap queue - 6664:280
1000万在int的取值范围内,不确定错误原因,相关代码如下:
#include <stdio.h> #include <stdlib.h> #include "mpi.h" #include <time.h> int main(int argc, char *argv[]){ int process_Rank, size_Of_Cluster; int number_of_elements; MPI_Init(&argc, &argv); MPI_Comm_size(MPI_COMM_WORLD, &size_Of_Cluster); MPI_Comm_rank(MPI_COMM_WORLD, &process_Rank); if(process_Rank == 0){ printf("Enter the number of elements:\n"); fflush(stdout); scanf("%d", &number_of_elements); } MPI_Bcast(&number_of_elements,1,MPI_INT, 0, MPI_COMM_WORLD); int *outputs = (int*)malloc(number_of_elements * sizeof(int)); unsigned long long chunk_size = number_of_elements/ size_Of_Cluster; int my_input[chunk_size], my_output[chunk_size]; for(int i = 0; i < number_of_elements; i++){ outputs[i] = i+1; } MPI_Barrier(MPI_COMM_WORLD); clock_t begin = clock(); MPI_Scatter(outputs, chunk_size, MPI_INT, &my_input, chunk_size, MPI_INT, 0, MPI_COMM_WORLD); for(int i = 0; i <= chunk_size; i++){ my_output[i] = my_input[i]; } MPI_Gather(&my_output, chunk_size, MPI_INT, outputs, chunk_size, MPI_INT, 0, MPI_COMM_WORLD); int iterate_terms[5] = {2,4,8,4,2}; int starting_terms[5] = {1,3,7,3,1}; int subtract_terms[5] = {1,2,4,0,0}; int adding_terms[5] = {0,0,0,2,1}; for(int j = 0; j < 5; j++){ MPI_Scatter(outputs, chunk_size, MPI_INT, &my_input, chunk_size, MPI_INT, 0, MPI_COMM_WORLD); for(int i = starting_terms[j]; i <= chunk_size; i+= iterate_terms[j]){ my_output[i+adding_terms[j]] += my_input[i-subtract_terms[j]]; } MPI_Gather(&my_output, chunk_size, MPI_INT, outputs, chunk_size, MPI_INT, 0, MPI_COMM_WORLD); } MPI_Barrier(MPI_COMM_WORLD); if(process_Rank == 0){ for(int i = chunk_size-1; i < number_of_elements; i+=chunk_size){ outputs[i+1] += outputs[i]; outputs[i+2] += outputs[i]; outputs[i+3] += outputs[i]; } clock_t end = clock(); double time_spent = (double)(end-begin) / CLOCKS_PER_SEC; for(int i = 0; i < number_of_elements; i++){ printf("%d \n", outputs[i]); fflush(stdout); } printf("took %f", time_spent); fflush(stdout); } else { clock_t end = clock(); } MPI_Finalize(); return 0; }
核心问题:栈溢出引发的进程崩溃
错误提示指向MPI_Bcast,但实际根源是栈上分配的变长数组(VLA)超出了进程默认栈大小限制。
代码中这两行是关键问题点:
unsigned long long chunk_size = number_of_elements/ size_Of_Cluster; int my_input[chunk_size], my_output[chunk_size];
当N=1000万时,4节点下chunk_size为250万,每个int占4字节,单个数组就需要10MB内存,两个数组合计20MB。而多数操作系统的默认栈大小仅为8-10MB,栈上分配这么大的数组会直接触发栈溢出,导致进程异常终止,进而引发MPI通信错误(进程崩溃后无法完成广播操作)。
修复方案
将栈上的变长数组改为堆上动态分配,避免栈溢出:
unsigned long long chunk_size = number_of_elements/ size_Of_Cluster; // 替换栈上数组为堆分配 int *my_input = (int*)malloc(chunk_size * sizeof(int)); int *my_output = (int*)malloc(chunk_size * sizeof(int)); // 程序结束前释放内存 if (my_input != NULL) free(my_input); if (my_output != NULL) free(my_output);
其他潜在问题修复
- 检查内存分配是否成功:动态分配后必须判断是否分配成功,避免空指针访问:
int *outputs = (int*)malloc(number_of_elements * sizeof(int)); if (outputs == NULL) { fprintf(stderr, "Process %d: 分配outputs数组失败\n", process_Rank); MPI_Abort(MPI_COMM_WORLD, 1); }
- 修复循环越界问题:代码中多处循环条件错误,比如:
for(int i = 0; i <= chunk_size; i++){ my_output[i] = my_input[i]; }
数组下标范围是0到chunk_size-1,应改为i < chunk_size,否则会访问越界内存。同理,后续循环中也要确保i+adding_terms[j]不超过chunk_size-1。
- 处理N无法被节点数整除的情况:当前代码假设N能被4整除,若不能,最后一个节点的chunk_size会更小,需改用
MPI_Scatterv和MPI_Gatherv处理不等长数据分发,避免通信出错。
总结
直接导致崩溃的原因是栈上大数组引发的栈溢出,改成堆分配即可解决核心问题。同时修复循环越界、内存分配检查等问题,能大幅提升程序健壮性。
内容的提问来源于stack exchange,提问作者Matthew Haywood

