MPI_Gatherv无法正常工作,求排查矩阵向量乘积并行代码问题
MPI_Gatherv无法收集非root进程数据的问题排查与修复
核心问题:ylocal数组索引越界
你的代码中,矩阵向量乘积计算时对ylocal的赋值使用了全局行号i作为索引,但ylocal是每个进程的本地向量,长度仅为local_rows(对应进程负责的行数)。这会导致内存越界,计算结果被写入到ylocal数组以外的内存区域,MPI_Gatherv读取的自然不是正确的计算数据。
以mpirun -np 3 ./mxvnm 10 10为例:
- 进程0负责行0-2(local_rows=3),i从0到2,
ylocal[i]是合法索引; - 进程1负责行3-5(local_rows=3),i从3到5,
ylocal[3]、ylocal[4]、ylocal[5]超出了ylocal的数组范围(仅0-2有效); - 进程2负责行6-9(local_rows=4),i从6到9,
ylocal[6]等同样越界。
修复步骤
1. 修正ylocal的赋值索引
为每个进程计算本地起始行号,循环时用本地索引访问ylocal:
// 新增:计算本地起始行 int start_row = myid * np; /* Matrix-vector product, y = Ax */ startMPI = MPI_Wtime(); for(i = start_row; i < limit; i++){ temp = 0.0; for(j=0; j<M; j++){ temp += A[i][j] * x[j]; } // 使用本地索引:i - start_row int local_idx = i - start_row; ylocal[local_idx] = temp; printf("PROC[%d]: despues ylocal[%d]=%g (全局行号%d)\n", myid, local_idx, ylocal[local_idx], i); }
2. 优化MPI_Gatherv的参数(可选但推荐)
非root进程无需提供counts和displacements参数,可直接传NULL,避免未初始化变量的潜在问题:
MPI_Gatherv(ylocal, local_rows, MPI_FLOAT, y, (myid == 0 ? counts : NULL), (myid == 0 ? displacements : NULL), MPI_FLOAT, 0, MPI_COMM_WORLD);
3. 可选优化:矩阵与向量的分布式初始化
当前所有进程都独立分配并初始化全局矩阵A和向量x,这会浪费内存(尤其是大矩阵)。建议仅在root进程分配初始化,再通过MPI_Bcast广播给其他进程:
// 仅root分配并初始化A和x if(myid == 0){ if((Avector = (float *) malloc(N*M*sizeof(float))) == NULL) printf("Error en malloc Avector[%d]\n",N*M); if((A = (float **) malloc(N*sizeof(float *))) == NULL) printf("Error en malloc del array de %d punteros\n",N); for(i=0;i<N;i++) *(A+i) = Avector+i*M; // 矩阵初始化 for(i=0; i<N; i++) for(j=0; j<M; j++) A[i][j] = (0.15*i - 0.1*j)/N; // 分配并初始化x if((x = (float *) malloc(M*sizeof(float))) == NULL) printf("Error en malloc x[%d]\n",M); for(i=0; i<M; i++) x[i] = (M/2.0 - i); } // 广播矩阵A MPI_Bcast(Avector, N*M, MPI_FLOAT, 0, MPI_COMM_WORLD); // 广播向量x MPI_Bcast(x, M, MPI_FLOAT, 0, MPI_COMM_WORLD); // 非root进程初始化A的指针数组 if(myid != 0){ if((A = (float **) malloc(N*sizeof(float *))) == NULL) printf("Error en malloc del array de %d punteros\n",N); for(i=0;i<N;i++) *(A+i) = Avector+i*M; }
修复后的验证
运行mpirun -np 3 ./mxvnm 10 10后,root进程输出的y数组应该能正确显示所有进程的计算结果,而非仅root进程的数据。
内容的提问来源于stack exchange,提问作者Nementh
相关产品推荐
相关产品推荐

