含内存填充的结构体使用MPI_Get性能异常缓慢的原因与解决办法
MPI单边通信与双边通信性能差异问题
我开发了一款分布式环境下传输std::vector内容的MPI应用,分别实现了基于MPI_Window+MPI_Get的单边通信方案,以及基于MPI_SendRecv的双边点对点通信方案。测试中发现,基于MPI_Get的实现性能比后者慢10至100倍,最终定位到问题出在传输包含内存填充的自定义结构体时,MPI_Get出现异常缓慢的情况。
以下是测试用例,分别用两种通信方式传输两种结构体:
Test1:{double, int},实际有效数据12字节,因内存对齐需要额外4字节填充,总大小16字节Test2:{double, int, int},有效数据16字节,无内存填充,总大小16字节
测试环境为Open MPI 4.1.2和Intel MPI 2021.6,两种环境下结果一致(Intel MPI中仅当进程数np>2时触发实际通信)。
测试代码
#include <mpi.h> #include <chrono> #include <iostream> #include <cassert> #include <type_traits> struct Test1 { double a; int b; }; bool operator!=(const Test1 x, const Test1 y) {return x.a != y.a or x.b != y.b;}; struct Test2 { double a; int b, c; }; bool operator!=(const Test2 x, const Test2 y) {return x.a != y.a or x.b != y.b or x.c != y.c;}; const int maxWorldSize = 8; //each proc sends an array of sz elements of type T to every other proc with SendRecv template<class T, int sz> double SendRecv(int world_size, int world_rank, MPI_Datatype dt) { T sendBuffer[sz], recvBuffer[maxWorldSize*sz]; T element; if constexpr (std::is_same_v<T, Test1>) element = {1., 2}; else if constexpr (std::is_same_v<T, Test2>) element = {1., 2, 3}; for (int i = 0; i < sz; i++) { sendBuffer[i] = element; } auto startTime = std::chrono::system_clock::now(); for (int offset = 1; offset < world_size; offset++) { int dest = (world_rank + offset)%world_size; int source = (world_rank - offset + world_size)%world_size; MPI_Status s; MPI_Sendrecv(sendBuffer, sz, dt, dest, 0, recvBuffer + source*sz, sz, dt, source, 0, MPI_COMM_WORLD, &s); } auto endTime = std::chrono::system_clock::now(); //ensure the result is used somewhere for (int i = 0; i < world_size; i++) { if (i != world_rank) { for (int j = 0; j < sz; j++) { if (recvBuffer[i*sz + j] != element) std::cout << "Problem" << std::endl; } } } return std::chrono::duration<double>(endTime - startTime).count(); } //each proc sends an array of sz elements of type T to every other proc with Window and Get template<class T, int sz> double Get(int world_size, int world_rank, MPI_Datatype dt) { T sendBuffer[sz], recvBuffer[maxWorldSize*sz]; T element; if constexpr (std::is_same_v<T, Test1>) element = {1., 2}; else if constexpr (std::is_same_v<T, Test2>) element = {1., 2, 3}; for (int i = 0; i < sz; i++) { sendBuffer[i] = element; } MPI_Win window; MPI_Info infos; MPI_Info_create(&infos); MPI_Win_create(sendBuffer, sz*sizeof(T), sizeof(T), infos, MPI_COMM_WORLD, &window); MPI_Win_fence(0, window); auto startTime = std::chrono::system_clock::now(); for (int offset = 1; offset < world_size; offset++) { int source = (world_rank - offset + world_size)%world_size; MPI_Get(recvBuffer + source*sz, sz, dt, source, 0, sz, dt, window); } auto endTime = std::chrono::system_clock::now(); MPI_Win_fence(0, window); MPI_Win_free(&window); MPI_Info_free(&infos); //ensure the result is used somewhere for (int i = 0; i < world_size; i++) { if (i != world_rank) { for (int j = 0; j < sz; j++) { if (recvBuffer[i*sz + j] != element) std::cout << "Problem" << std::endl; } } } return std::chrono::duration<double>(endTime - startTime).count(); } int main(int argc, char **argv) { MPI_Init(&argc, &argv); int world_size, world_rank; MPI_Comm_size(MPI_COMM_WORLD, &world_size); MPI_Comm_rank(MPI_COMM_WORLD, &world_rank); if (world_size > maxWorldSize) std::cout << "SendRecv and Get function will seg fault if world_size > " << maxWorldSize << std::endl; if (world_size == 1) std::cout << "MPI_Window won't work if world_size == 1" << std::endl; assert(world_size <= maxWorldSize && world_size > 1); const int size = 1e6; MPI_Datatype datatype1; int array_of_blocklengths1[2] = {1, 1}; MPI_Aint array_of_displacements1[2] = {0, 8}; MPI_Datatype array_of_types1[2] = {MPI_DOUBLE, MPI_INT}; MPI_Type_create_struct(2, array_of_blocklengths1, array_of_displacements1, array_of_types1, &datatype1); MPI_Type_commit(&datatype1); //send {double, int} with MPI_SendRecv auto time = SendRecv<Test1, size>(world_size, world_rank, datatype1); if (world_rank == 0) std::cout << "{double, int} SendRecv : " << time << " s" << std::endl; //send {double, int} with MPI_Window and MPI_Get time = Get<Test1, size>(world_size, world_rank, datatype1); if (world_rank == 0) std::cout << "{double, int} Get : " << time << " s" << std::endl; MPI_Datatype datatype2; int array_of_blocklengths2[2] = {1, 2}; MPI_Aint array_of_displacements2[2] = {0, 8}; MPI_Datatype array_of_types2[2] = {MPI_DOUBLE, MPI_INT}; MPI_Type_create_struct(2, array_of_blocklengths2, array_of_displacements2, array_of_types2, &datatype2); MPI_Type_commit(&datatype2); //send {double, int, int} with MPI_SendRecv time = SendRecv<Test1, size>(world_size, world_rank, datatype2); if (world_rank == 0) std::cout << "{double, int, int} SendRecv : " << time << " s" << std::endl; //send {double, int, int} with MPI_Window and MPI_Get time = Get<Test1, size>(world_size, world_rank, datatype2); if (world_rank == 0) std::cout << "{double, int, int} Get : " << time << " s" << std::endl; MPI_Finalize(); }
测试结果
mpirun -np 4 ./compareGetWithSendRecv {double, int} SendRecv : 0.0303547 s {double, int} Get : 1.9196 s {double, int, int} SendRecv : 0.0164659 s {double, int, int} Get : 0.0147757 s
问题解答
1. 该结果是否正常?
是正常的,核心原因在于两种通信方式对非连续内存布局的处理逻辑不同:
MPI_SendRecv属于双边通信,发送端会根据自定义MPI类型,将结构体中的有效数据打包成连续的缓冲区后传输,接收端再解包还原,全程可以利用连续内存的批量传输优化(如DMA),因此效率不受结构体填充影响。MPI_Get属于单边通信,直接从源进程的内存中读取数据。当结构体存在填充时,自定义MPI类型描述的是非连续的内存块(每个元素仅包含12字节有效数据,间隔4字节填充),MPI无法将多个元素合并为连续的内存区域读取,只能逐个处理非连续块,导致大量小数据传输,开销剧增。而无填充的Test2结构体内存布局连续,MPI_Get可以直接批量读取,性能与MPI_SendRecv接近。
2. 除添加冗余字段外的解决方案
(1)调整自定义MPI类型的Extent
通过MPI_Type_create_resized修改自定义类型的跨度(Extent),使其匹配结构体的实际内存大小(包含填充),让MPI识别出内存中连续的元素块,从而启用批量传输优化。
示例修改datatype1的创建逻辑:
MPI_Datatype datatype1_temp; int array_of_blocklengths1[2] = {1, 1}; MPI_Aint array_of_displacements1[2] = {0, 8}; MPI_Datatype array_of_types1[2] = {MPI_DOUBLE, MPI_INT}; MPI_Type_create_struct(2, array_of_blocklengths1, array_of_displacements1, array_of_types1, &datatype1_temp); // 将类型的Extent设置为结构体实际大小(16字节) MPI_Type_create_resized(datatype1_temp, 0, sizeof(Test1), &datatype1); MPI_Type_commit(&datatype1); MPI_Type_free(&datatype1_temp);
(2)使用编译器指令控制结构体填充
通过编译器特性强制结构体无填充(如GCC的__attribute__((packed))、MSVC的#pragma pack(push, 1)),让结构体内存布局连续。但需注意:未对齐的内存访问在部分架构(如ARM)上会导致崩溃,在x86架构上也可能带来性能损耗,需权衡使用。
示例:
struct Test1 __attribute__((packed)) { double a; int b; };
(3)转换为连续内存缓冲区传输
将结构体的有效数据复制到连续的数组中(如分离double和int到两个独立数组,或用字节数组打包),传输完成后再解析回结构体。这种方式完全规避了内存填充的影响,适用于所有场景。
内容的提问来源于stack exchange,提问作者Antoine Motte
相关产品推荐
相关产品推荐

