MPI笛卡尔拓扑移除主进程实现矩阵乘法及解决Cart_shift报错
问题分析与解决方案
你的核心问题是主进程(rank=num_procs-1)不属于创建的子通信子,却执行了MPI_Cart_create等拓扑操作,导致MPI_Cart_shift报错。另外原代码还有几个细节问题需要修正:
关键错误点
- 主进程不在
new_group中,MPI_Comm_create返回的old_comm_2d对它来说是MPI_COMM_NULL,此时调用MPI_Cart_create会直接触发错误。 sqrt返回浮点型,直接赋值给int类型的dims可能存在精度丢失(比如当num_procs-1是完全平方数时,浮点计算可能得到p-0.999999,转int后变成p-1)。- 未释放malloc分配的内存,存在内存泄漏。
修正后的代码
#include <stdio.h> #include <stdlib.h> #include <time.h> #include <math.h> #include <mpi.h> void self_checking(int N, int num_procs) { MPI_Status status; int world_rank; MPI_Comm_rank(MPI_COMM_WORLD, &world_rank); MPI_Group world_group; MPI_Comm_group(MPI_COMM_WORLD, &world_group); int *ranks = (int *)malloc((num_procs - 1) * sizeof(int)); for(int i = 0; i < num_procs - 1; i++) { ranks[i] = i; } MPI_Group new_group; MPI_Group_incl(world_group, num_procs - 1, ranks, &new_group); MPI_Comm old_comm_2d = MPI_COMM_NULL; MPI_Comm comm_2d = MPI_COMM_NULL; // 只有非主进程参与子通信子创建与拓扑初始化 if(world_rank != num_procs - 1) { MPI_Comm_create(MPI_COMM_WORLD, new_group, &old_comm_2d); int *dims = (int *)malloc(2 * sizeof(int)); int *periods = (int *)malloc(2 * sizeof(int)); // 强制转换为int,确保维度正确 int grid_size = (int)sqrt(num_procs - 1); // 校验是否为完全平方数,避免拓扑创建失败 if(grid_size * grid_size != num_procs - 1) { if(world_rank == 0) fprintf(stderr, "Error: num_procs-1 must be a perfect square!\n"); MPI_Abort(MPI_COMM_WORLD, 1); } dims[0] = dims[1] = grid_size; periods[0] = periods[1] = 1; MPI_Cart_create(old_comm_2d, 2, dims, periods, 1, &comm_2d); int local_rank; MPI_Comm_rank(comm_2d, &local_rank); int *coords = (int *)malloc(2 * sizeof(int)); MPI_Cart_coords(comm_2d, local_rank, 2, coords); srand((unsigned int)time(NULL) + local_rank); int leftrank, rightrank, uprank, downrank; MPI_Cart_shift(comm_2d, 0, -1, &downrank, &uprank); MPI_Cart_shift(comm_2d, 1, -1, &rightrank, &leftrank); printf("World rank %d, Local rank %d, coords (%d,%d), left %d, right %d, up %d, down %d\n", world_rank, local_rank, coords[0], coords[1], leftrank, rightrank, uprank, downrank); // 释放局部内存 free(coords); free(dims); free(periods); } else { printf("Main process, world rank %d\n", world_rank); // 主进程在这里执行结果校验逻辑 } // 释放MPI资源与全局内存 if(old_comm_2d != MPI_COMM_NULL) MPI_Comm_free(&old_comm_2d); if(comm_2d != MPI_COMM_NULL) MPI_Comm_free(&comm_2d); MPI_Group_free(&new_group); MPI_Group_free(&world_group); free(ranks); }
核心修改说明
- 隔离主进程逻辑:将子通信子创建、拓扑初始化等操作放在非主进程的分支中,主进程只执行结果校验逻辑,完全不参与拓扑相关调用。
- 修复维度计算:将
sqrt结果强制转为int,并添加完全平方数校验,避免因浮点精度问题导致拓扑维度错误。 - 资源释放:所有malloc分配的内存、MPI通信子和组资源都进行了释放,避免内存泄漏和资源占用。
- 区分rank变量:用
world_rank表示全局rank,local_rank表示拓扑通信子内的rank,避免变量混淆。
内容的提问来源于stack exchange,提问作者Delev1n
相关产品推荐
相关产品推荐

