MPI程序使用MPI_Wtime()计算总时间时出错,求技术支持
MPI积分程序报错MPI_ERR_RANK的修复方案
问题背景
尝试统计每个处理器的计算耗时及程序总耗时,但运行时触发MPI_ERR_RANK错误,此前相同计时逻辑在其他代码中可正常运行。
报错信息
[Sid-Laptop:4987] *** An error occurred in MPI_Recv [Sid-Laptop:4987] *** reported by process [852688897,2] [Sid-Laptop:4987] *** on communicator MPI_COMM_WORLD [Sid-Laptop:4987] *** MPI_ERR_RANK: invalid rank [Sid-Laptop:4987] *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, [Sid-Laptop:4987] *** and potentially your MPI job) Enter a, b and n [Sid-Laptop:04980] 2 more processes have sent help message help-mpi-errors.txt / mpi_errors_are_fatal [Sid-Laptop:04980] Set MCA parameter "orte_base_help_aggregate" to 0 to see all help / error messages
问题分析
- 未初始化的source变量:非0进程执行
MPI_Recv时,source变量未赋值,导致MPI尝试从非法进程号接收数据,触发MPI_ERR_RANK。数据由进程0发送,需将source固定为0。 - Trap函数缺失返回值:
Trap函数计算完积分后无返回语句,导致integral变量值未定义,后续通信和计算出现异常。 - local_b计算逻辑错误:原代码
local_b = (local_a + local_n) * h逻辑错误,正确区间终点应为local_a + local_n * h。 - n无法被进程数整除的场景遗漏:当前
local_n = n/size只处理了n能被进程数整除的情况,若无法整除会导致部分区间未被计算。
修复后的完整代码
#include <stdio.h> #include <stdlib.h> #include <time.h> #include "mpi.h" float Trap(float local_a, float local_b, int local_n, float h); float f(float x); int main(int argc, char** argv){ int my_rank; double time1, time2, duration, global; int size; float a ; float b ; int n ; float h; float local_a; float local_b; int local_n; float integral; float total; int source = 0; // 初始化source为进程0 int dest = 0; int tag = 0; MPI_Status status; MPI_Init(&argc, &argv); MPI_Comm_rank(MPI_COMM_WORLD, &my_rank); MPI_Comm_size(MPI_COMM_WORLD, &size); if (my_rank == 0){ printf("Enter a, b and n\n"); scanf("%f %f %d", &a, &b, &n); for ( dest = 1 ; dest < size; dest++){ MPI_Send(&a, 1 , MPI_FLOAT, dest , 0, MPI_COMM_WORLD); MPI_Send(&b, 1 , MPI_FLOAT, dest , 1, MPI_COMM_WORLD); MPI_Send(&n, 1 , MPI_INT, dest , 2, MPI_COMM_WORLD); } } else{ // 从进程0接收数据 MPI_Recv(&a, 1, MPI_FLOAT, source, 0, MPI_COMM_WORLD, &status); MPI_Recv(&b, 1, MPI_FLOAT, source, 1, MPI_COMM_WORLD, &status); MPI_Recv(&n, 1, MPI_INT, source, 2, MPI_COMM_WORLD, &status); } MPI_Barrier(MPI_COMM_WORLD); time1 = MPI_Wtime(); h = (b - a)/n; local_n = n / size; // 处理n无法被size整除的情况,最后一个进程承担剩余任务 if(my_rank == size - 1){ local_n += n % size; } local_a = a + my_rank * (n / size) * h; // 修正区间终点计算逻辑 local_b = local_a + local_n * h; integral = Trap(local_a, local_b, local_n, h); if (my_rank == 0){ total = integral; for (source = 1; source < size; source++){ MPI_Recv(&integral, 1, MPI_FLOAT, source, tag, MPI_COMM_WORLD, &status); total += integral; } } else { MPI_Send(&integral, 1, MPI_FLOAT, dest, tag, MPI_COMM_WORLD); } time2 = MPI_Wtime(); duration = time2 - time1; MPI_Reduce(&duration, &global,1,MPI_DOUBLE,MPI_SUM,0,MPI_COMM_WORLD); if (my_rank == 0){ printf("With n = %d trapezoids, our estimate\n", n); printf("of the integral from %f to %f = %0.8f\n",a,b,total); printf("Global runtime is %f\n",global); } printf("Runtime at %d is %f\n", my_rank,duration); MPI_Finalize(); return 0; } // 添加返回语句,返回计算好的积分值 float Trap(float local_a, float local_b, int local_n, float h){ float integral; float x; integral = (f(local_a) + f(local_b))/2.0; x = local_a; for (int i = 1; i <= local_n-1; i++){ x += h; integral += f(x); } integral *= h; return integral; } float f(float x){ return x*x; }
修复说明
- 初始化
source为0,确保非0进程从正确节点接收数据; - 给
Trap函数添加返回语句,传递计算完成的积分值; - 修正
local_b的计算逻辑,保证区间划分正确; - 补充n无法被进程数整除时的处理逻辑,避免遗漏计算区间;
- 主函数末尾添加
return 0;,符合C语言规范。
内容的提问来源于stack exchange,提问作者h3avyc0der
相关产品推荐
相关产品推荐

