You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MPI_Gatherv无法正常工作,求排查矩阵向量乘积并行代码问题

MPI_Gatherv无法收集非root进程数据的问题排查与修复

核心问题:ylocal数组索引越界

你的代码中,矩阵向量乘积计算时对ylocal的赋值使用了全局行号i作为索引,但ylocal是每个进程的本地向量,长度仅为local_rows(对应进程负责的行数)。这会导致内存越界,计算结果被写入到ylocal数组以外的内存区域,MPI_Gatherv读取的自然不是正确的计算数据。

以mpirun -np 3 ./mxvnm 10 10为例:

  • 进程0负责行0-2(local_rows=3),i从0到2,ylocal[i]是合法索引;
  • 进程1负责行3-5(local_rows=3),i从3到5,ylocal[3]、ylocal[4]、ylocal[5]超出了ylocal的数组范围(仅0-2有效);
  • 进程2负责行6-9(local_rows=4),i从6到9,ylocal[6]等同样越界。

修复步骤

1. 修正ylocal的赋值索引

为每个进程计算本地起始行号,循环时用本地索引访问ylocal:

// 新增:计算本地起始行
int start_row = myid * np;

/* Matrix-vector product, y = Ax */    
startMPI = MPI_Wtime();
for(i = start_row; i < limit; i++){
    temp = 0.0;
    for(j=0; j<M; j++){
        temp += A[i][j] * x[j];     
    }        
    // 使用本地索引:i - start_row
    int local_idx = i - start_row;
    ylocal[local_idx] = temp;
    printf("PROC[%d]: despues ylocal[%d]=%g (全局行号%d)\n", myid, local_idx, ylocal[local_idx], i);
}

2. 优化MPI_Gatherv的参数(可选但推荐)

非root进程无需提供counts和displacements参数,可直接传NULL,避免未初始化变量的潜在问题:

MPI_Gatherv(ylocal, local_rows, MPI_FLOAT, 
            y, (myid == 0 ? counts : NULL), (myid == 0 ? displacements : NULL), 
            MPI_FLOAT, 0, MPI_COMM_WORLD);

3. 可选优化:矩阵与向量的分布式初始化

当前所有进程都独立分配并初始化全局矩阵A和向量x,这会浪费内存(尤其是大矩阵)。建议仅在root进程分配初始化,再通过MPI_Bcast广播给其他进程:

// 仅root分配并初始化A和x
if(myid == 0){
    if((Avector = (float *) malloc(N*M*sizeof(float))) == NULL)
        printf("Error en malloc Avector[%d]\n",N*M);
    if((A = (float **) malloc(N*sizeof(float *))) == NULL)
        printf("Error en malloc del array de %d punteros\n",N);
    for(i=0;i<N;i++)
        *(A+i) = Avector+i*M;
    // 矩阵初始化
    for(i=0; i<N; i++)
        for(j=0; j<M; j++)
            A[i][j] = (0.15*i - 0.1*j)/N;
    // 分配并初始化x
    if((x = (float *) malloc(M*sizeof(float))) == NULL)
        printf("Error en malloc x[%d]\n",M);
    for(i=0; i<M; i++)
        x[i] = (M/2.0 - i);
}
// 广播矩阵A
MPI_Bcast(Avector, N*M, MPI_FLOAT, 0, MPI_COMM_WORLD);
// 广播向量x
MPI_Bcast(x, M, MPI_FLOAT, 0, MPI_COMM_WORLD);
// 非root进程初始化A的指针数组
if(myid != 0){
    if((A = (float **) malloc(N*sizeof(float *))) == NULL)
        printf("Error en malloc del array de %d punteros\n",N);
    for(i=0;i<N;i++)
        *(A+i) = Avector+i*M;
}

修复后的验证

运行mpirun -np 3 ./mxvnm 10 10后,root进程输出的y数组应该能正确显示所有进程的计算结果,而非仅root进程的数据。

内容的提问来源于stack exchange,提问作者Nementh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 08:25:04