You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MPI笛卡尔拓扑移除主进程实现矩阵乘法及解决Cart_shift报错

问题分析与解决方案

你的核心问题是主进程(rank=num_procs-1)不属于创建的子通信子,却执行了MPI_Cart_create等拓扑操作,导致MPI_Cart_shift报错。另外原代码还有几个细节问题需要修正:

关键错误点

  1. 主进程不在new_group中,MPI_Comm_create返回的old_comm_2d对它来说是MPI_COMM_NULL,此时调用MPI_Cart_create会直接触发错误。
  2. sqrt返回浮点型,直接赋值给int类型的dims可能存在精度丢失(比如当num_procs-1是完全平方数时,浮点计算可能得到p-0.999999,转int后变成p-1)。
  3. 未释放malloc分配的内存,存在内存泄漏。

修正后的代码

#include <stdio.h>
#include <stdlib.h>
#include <time.h>
#include <math.h>
#include <mpi.h>

void self_checking(int N, int num_procs)
{
    MPI_Status status;
    int world_rank;
    MPI_Comm_rank(MPI_COMM_WORLD, &world_rank);

    MPI_Group world_group;
    MPI_Comm_group(MPI_COMM_WORLD, &world_group);
 
    int *ranks = (int *)malloc((num_procs - 1) * sizeof(int));
    for(int i = 0; i < num_procs - 1; i++)
    {
        ranks[i] = i;
    }
    MPI_Group new_group;
    MPI_Group_incl(world_group, num_procs - 1, ranks, &new_group);
 
    MPI_Comm old_comm_2d = MPI_COMM_NULL;
    MPI_Comm comm_2d = MPI_COMM_NULL;

    // 只有非主进程参与子通信子创建与拓扑初始化
    if(world_rank != num_procs - 1)
    {
        MPI_Comm_create(MPI_COMM_WORLD, new_group, &old_comm_2d);

        int *dims = (int *)malloc(2 * sizeof(int));
        int *periods = (int *)malloc(2 * sizeof(int));
        // 强制转换为int,确保维度正确
        int grid_size = (int)sqrt(num_procs - 1);
        // 校验是否为完全平方数,避免拓扑创建失败
        if(grid_size * grid_size != num_procs - 1)
        {
            if(world_rank == 0)
                fprintf(stderr, "Error: num_procs-1 must be a perfect square!\n");
            MPI_Abort(MPI_COMM_WORLD, 1);
        }
        dims[0] = dims[1] = grid_size;
        periods[0] = periods[1] = 1;
        MPI_Cart_create(old_comm_2d, 2, dims, periods, 1, &comm_2d);

        int local_rank;
        MPI_Comm_rank(comm_2d, &local_rank);
        int *coords = (int *)malloc(2 * sizeof(int));
        MPI_Cart_coords(comm_2d, local_rank, 2, coords);

        srand((unsigned int)time(NULL) + local_rank);

        int leftrank, rightrank, uprank, downrank;
        MPI_Cart_shift(comm_2d, 0, -1, &downrank, &uprank);
        MPI_Cart_shift(comm_2d, 1, -1, &rightrank, &leftrank);

        printf("World rank %d, Local rank %d, coords (%d,%d), left %d, right %d, up %d, down %d\n", 
               world_rank, local_rank, coords[0], coords[1], leftrank, rightrank, uprank, downrank);

        // 释放局部内存
        free(coords);
        free(dims);
        free(periods);
    }
    else
    {
        printf("Main process, world rank %d\n", world_rank);
        // 主进程在这里执行结果校验逻辑
    }

    // 释放MPI资源与全局内存
    if(old_comm_2d != MPI_COMM_NULL)
        MPI_Comm_free(&old_comm_2d);
    if(comm_2d != MPI_COMM_NULL)
        MPI_Comm_free(&comm_2d);
    MPI_Group_free(&new_group);
    MPI_Group_free(&world_group);
    free(ranks);
}

核心修改说明

  • 隔离主进程逻辑:将子通信子创建、拓扑初始化等操作放在非主进程的分支中,主进程只执行结果校验逻辑,完全不参与拓扑相关调用。
  • 修复维度计算:将sqrt结果强制转为int,并添加完全平方数校验,避免因浮点精度问题导致拓扑维度错误。
  • 资源释放:所有malloc分配的内存、MPI通信子和组资源都进行了释放,避免内存泄漏和资源占用。
  • 区分rank变量:用world_rank表示全局rank,local_rank表示拓扑通信子内的rank,避免变量混淆。

内容的提问来源于stack exchange,提问作者Delev1n

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 06:23:27