C语言MPI多进程求和代码进程越多运行越慢问题排查
问题描述
我正在学习基于C语言的MPI,编写了一段对包含1e8个元素的向量求和的代码。编译命令为:mpicc -o vector_sum vector_send.c -lm,运行命令为time mpirun -n x vector_sum。测试发现x=2时的运行速度比x=5时更快——进程越多,运行时间反而越长。我原本认为各进程的求和操作相互独立,不该出现性能下降的情况,怀疑是进程间数据发送的循环导致的问题,想确认代码是否存在错误。
原代码
/* Code which splits a vector and send information to other processes. In case of main vector does not split equally to all processes, the leftover is passed to process id 1. Process id 0 is the root process. Therefore it does not count while passing information. Each process will calculate the partial sum of vector values and send it back to root process, which will calculate the total sum. Since the processes are independent, the printing order will be different at each run. compile as: mpicc -o vector_sum vector_send.c -lm run as: time mpirun -n x vector_sum x = number of splits desired + root process. For example: if x = 3, the vector will be splited in two. */ #include<stdio.h> #include<mpi.h> #include<math.h> #define vec_len 100000000 double vec1[vec_len]; double vec2[vec_len]; int main(int argc, char* argv[]){ // defining program variables int i; double sum, partial_sum; // defining parallel step variables int my_id, num_proc, ierr, an_id, root_process; // id of process and total number of processes int num_2_send, num_2_recv, start_point, vec_size, rows_per_proc, leftover; ierr = MPI_Init(&argc, &argv); root_process = 0; ierr = MPI_Comm_size(MPI_COMM_WORLD, &num_proc); ierr = MPI_Comm_rank(MPI_COMM_WORLD, &my_id); if(my_id == root_process){ // Root process: Define vector size, how to split vector and send information to workers vec_size = 1e8; // size of main vector for(i = 0; i < vec_size; i++){ //vec1[i] = pow(-1.0,i+2)/(2.0*(i+1)-1.0); // defining main vector... Correct answer for total sum = 0.78539816339 vec1[i] = pow(i,2)+1.0; // defining main vector... //printf("Main vector position %d: %f\n", i, vec1[i]); // uncomment if you wish to print the main vector } rows_per_proc = vec_size / (num_proc - 1); // average values per process: using (num_proc - 1) because proc 0 does not count as a worker. rows_per_proc = floor(rows_per_proc); // getting the maximum integer possible. leftover = vec_size - (num_proc - 1)*rows_per_proc; // counting the leftover. // spliting and sending the values for(an_id = 1; an_id < num_proc; an_id++){ if(an_id == 1){ // worker id 1 will have more values if there is any leftover. num_2_send = rows_per_proc + leftover; // counting the amount of data to be sent. start_point = (an_id - 1)*num_2_send; // defining initial position in the main vector (data will be sent from here) } else{ num_2_send = rows_per_proc; start_point = (an_id - 1)*num_2_send + leftover; // starting point for other processes if there is leftover. } ierr = MPI_Send(&num_2_send, 1, MPI_INT, an_id, 1234, MPI_COMM_WORLD); // sending the information of how many data is going to workers. ierr = MPI_Send(&vec1[start_point], num_2_send, MPI_DOUBLE, an_id, 1234, MPI_COMM_WORLD); // sending pieces of the main vector. } sum = 0; for(an_id = 1; an_id < num_proc; an_id++){ ierr = MPI_Recv(&partial_sum, 1, MPI_DOUBLE, an_id, 4321, MPI_COMM_WORLD, MPI_STATUS_IGNORE); // recieving partial sum. sum = sum + partial_sum; } printf("Total sum = %f.\n", sum); } else{ // Workers:define which operation will be carried out by each one ierr = MPI_Recv(&num_2_recv, 1, MPI_INT, root_process, 1234, MPI_COMM_WORLD, MPI_STATUS_IGNORE); // recieving the information of how many data worker must expect. ierr = MPI_Recv(&vec2, num_2_recv, MPI_DOUBLE, root_process, 1234, MPI_COMM_WORLD, MPI_STATUS_IGNORE); // recieving main vector pieces. partial_sum = 0; for(i=0; i < num_2_recv; i++){ //printf("Position %d from worker id %d: %d\n", i, my_id, vec2[i]); // uncomment if you wish to print position, id and value of splitted vector partial_sum = partial_sum + vec2[i]; } printf("Partial sum of %d: %f\n",my_id, partial_sum); ierr = MPI_Send(&partial_sum, 1, MPI_DOUBLE, root_process, 4321, MPI_COMM_WORLD); // sending partial sum to root process. } ierr = MPI_Finalize(); }
问题根源分析
1. 串行数据分发的开销
根进程用循环串行调用MPI_Send给每个工作进程发数据,必须等待前一个发送完成才能处理下一个。进程越多,累计的发送等待和通信开销越大,直接拖慢整体速度。
2. 全局大数组的内存浪费
代码中定义了两个全局数组vec1和vec2,每个数组占800MB(1e8个double,每个8字节)。每个进程都会分配这两个数组,5个进程的总内存占用会超过4GB,极易触发内存分页/交换,导致性能暴跌。
3. 低效的向量初始化
用pow(i,2)计算平方是浮点库函数,远慢于直接的整数乘法(double)i*i,会增加初始化阶段的耗时。
4. 根进程闲置浪费
当前根进程只负责分发和汇总,不参与求和计算,没有充分利用计算资源。
优化修正方案
1. 用集体通信替代串行发送
使用MPI_Scatterv一次性完成数据分发,它专门用于向不同进程发送不同长度的数据块,效率远高于循环MPI_Send。
2. 动态分配内存
取消全局大数组,改用malloc为每个进程分配仅需的内存,大幅降低内存占用:
// 根进程分配主向量 double *vec1 = malloc(vec_size * sizeof(double)); // 工作进程根据接收长度分配局部向量 double *vec2 = malloc(num_2_recv * sizeof(double));
记得在使用完后调用free释放内存。
3. 优化初始化计算
将pow(i,2)替换为(double)i*i,初始化速度可提升数倍。
4. 让根进程参与计算
调整数据拆分逻辑,让根进程也处理一部分向量元素,充分利用所有进程的计算能力。
修正后的核心代码片段
int main(int argc, char* argv[]){ int my_id, num_proc; double local_sum = 0; MPI_Init(&argc, &argv); MPI_Comm_size(MPI_COMM_WORLD, &num_proc); MPI_Comm_rank(MPI_COMM_WORLD, &my_id); const int vec_size = 1e8; int rows_per_proc = vec_size / num_proc; int leftover = vec_size % num_proc; // 计算每个进程的处理长度和起始偏移 int sendcounts[num_proc]; int displs[num_proc]; int offset = 0; for(int i=0; i<num_proc; i++){ sendcounts[i] = rows_per_proc + (i < leftover ? 1 : 0); displs[i] = offset; offset += sendcounts[i]; } // 动态分配局部向量 double *local_vec = malloc(sendcounts[my_id] * sizeof(double)); // 根进程初始化主向量 double *vec1 = NULL; if(my_id == 0){ vec1 = malloc(vec_size * sizeof(double)); for(int i=0; i<vec_size; i++){ vec1[i] = (double)i*i + 1.0; } } // 集体通信分发数据 MPI_Scatterv(vec1, sendcounts, displs, MPI_DOUBLE, local_vec, sendcounts[my_id], MPI_DOUBLE, 0, MPI_COMM_WORLD); // 每个进程计算局部和 for(int i=0; i<sendcounts[my_id]; i++){ local_sum += local_vec[i]; } // 集体通信汇总结果 double total_sum; MPI_Reduce(&local_sum, &total_sum, 1, MPI_DOUBLE, MPI_SUM, 0, MPI_COMM_WORLD); if(my_id == 0){ printf("Total sum = %f.\n", total_sum); free(vec1); } free(local_vec); MPI_Finalize(); return 0; }
内容的提问来源于stack exchange,提问作者Felipe_SC

