You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

含内存填充的结构体使用MPI_Get性能异常缓慢的原因与解决办法

MPI单边通信与双边通信性能差异问题

我开发了一款分布式环境下传输std::vector内容的MPI应用,分别实现了基于MPI_Window+MPI_Get的单边通信方案,以及基于MPI_SendRecv的双边点对点通信方案。测试中发现,基于MPI_Get的实现性能比后者慢10至100倍,最终定位到问题出在传输包含内存填充的自定义结构体时,MPI_Get出现异常缓慢的情况。

以下是测试用例,分别用两种通信方式传输两种结构体:

  • Test1:{double, int},实际有效数据12字节,因内存对齐需要额外4字节填充,总大小16字节
  • Test2:{double, int, int},有效数据16字节,无内存填充,总大小16字节

测试环境为Open MPI 4.1.2和Intel MPI 2021.6,两种环境下结果一致(Intel MPI中仅当进程数np>2时触发实际通信)。

测试代码

#include <mpi.h>
#include <chrono>
#include <iostream>
#include <cassert>
#include <type_traits>

struct Test1
{
    double a;
    int b;
};

bool operator!=(const Test1 x, const Test1 y) {return x.a != y.a or x.b != y.b;};

struct Test2
{
    double a;
    int b, c;
};

bool operator!=(const Test2 x, const Test2 y) {return x.a != y.a or x.b != y.b or x.c != y.c;};

const int maxWorldSize = 8;

//each proc sends an array of sz elements of type T to every other proc with SendRecv
template<class T, int sz>
double SendRecv(int world_size, int world_rank, MPI_Datatype dt)
{
    T sendBuffer[sz], recvBuffer[maxWorldSize*sz];

    T element;
    if constexpr (std::is_same_v<T, Test1>) element = {1., 2};
    else if constexpr (std::is_same_v<T, Test2>) element = {1., 2, 3};

    for (int i = 0; i < sz; i++)
    {
        sendBuffer[i] = element;
    }

    auto startTime = std::chrono::system_clock::now();
    for (int offset = 1; offset < world_size; offset++)
    {
        int dest = (world_rank + offset)%world_size;
        int source = (world_rank - offset + world_size)%world_size;
        MPI_Status s;

        MPI_Sendrecv(sendBuffer, sz, dt, dest, 0, recvBuffer + source*sz, sz, dt, source, 0, MPI_COMM_WORLD, &s);
    }
    auto endTime = std::chrono::system_clock::now();

    //ensure the result is used somewhere
    for (int i = 0; i < world_size; i++)
    {
        if (i != world_rank)
        {
            for (int j = 0; j < sz; j++)
            {
                if (recvBuffer[i*sz + j] != element) std::cout << "Problem" << std::endl;
            }
        } 
    }

    return std::chrono::duration<double>(endTime - startTime).count();
}

//each proc sends an array of sz elements of type T to every other proc with Window and Get
template<class T, int sz>
double Get(int world_size, int world_rank, MPI_Datatype dt)
{
    T sendBuffer[sz], recvBuffer[maxWorldSize*sz];

    T element;
    if constexpr (std::is_same_v<T, Test1>) element = {1., 2};
    else if constexpr (std::is_same_v<T, Test2>) element = {1., 2, 3};

    for (int i = 0; i < sz; i++)
    {
        sendBuffer[i] = element;
    }

    MPI_Win window;
    MPI_Info infos;

    MPI_Info_create(&infos);
    MPI_Win_create(sendBuffer, sz*sizeof(T), sizeof(T), infos, MPI_COMM_WORLD, &window);
    MPI_Win_fence(0, window);

    auto startTime = std::chrono::system_clock::now();
    for (int offset = 1; offset < world_size; offset++)
    {
        int source = (world_rank - offset + world_size)%world_size;
        MPI_Get(recvBuffer + source*sz, sz, dt, source, 0, sz, dt, window);
    }
    auto endTime = std::chrono::system_clock::now();

    MPI_Win_fence(0, window);
    MPI_Win_free(&window);
    MPI_Info_free(&infos);

    //ensure the result is used somewhere
    for (int i = 0; i < world_size; i++)
    {
        if (i != world_rank)
        {
            for (int j = 0; j < sz; j++)
            {
                if (recvBuffer[i*sz + j] != element) std::cout << "Problem" << std::endl;
            }
        } 
    }

    return std::chrono::duration<double>(endTime - startTime).count();
}

int main(int argc, char **argv)
{
    MPI_Init(&argc, &argv);

    int world_size, world_rank;
    MPI_Comm_size(MPI_COMM_WORLD, &world_size);
    MPI_Comm_rank(MPI_COMM_WORLD, &world_rank);

    if (world_size > maxWorldSize) std::cout << "SendRecv and Get function will seg fault if world_size > " << maxWorldSize << std::endl;
    if (world_size == 1) std::cout << "MPI_Window won't work if world_size == 1" << std::endl;
    assert(world_size <= maxWorldSize && world_size > 1);

    const int size = 1e6;

    MPI_Datatype datatype1;
    int array_of_blocklengths1[2] = {1, 1};
    MPI_Aint array_of_displacements1[2] = {0, 8};
    MPI_Datatype array_of_types1[2] = {MPI_DOUBLE, MPI_INT};
    MPI_Type_create_struct(2, array_of_blocklengths1, array_of_displacements1, array_of_types1, &datatype1);
    MPI_Type_commit(&datatype1);
    
    //send {double, int} with MPI_SendRecv
    auto time = SendRecv<Test1, size>(world_size, world_rank, datatype1);
    if (world_rank == 0) std::cout << "{double, int} SendRecv : " << time << " s" << std::endl;

    //send {double, int} with MPI_Window and MPI_Get
    time = Get<Test1, size>(world_size, world_rank, datatype1);
    if (world_rank == 0) std::cout << "{double, int} Get : " << time << " s" << std::endl;

    MPI_Datatype datatype2;
    int array_of_blocklengths2[2] = {1, 2};
    MPI_Aint array_of_displacements2[2] = {0, 8};
    MPI_Datatype array_of_types2[2] = {MPI_DOUBLE, MPI_INT};
    MPI_Type_create_struct(2, array_of_blocklengths2, array_of_displacements2, array_of_types2, &datatype2);
    MPI_Type_commit(&datatype2);
    
    //send {double, int, int} with MPI_SendRecv
    time = SendRecv<Test1, size>(world_size, world_rank, datatype2);
    if (world_rank == 0) std::cout << "{double, int, int} SendRecv : " << time << " s" << std::endl;

    //send {double, int, int} with MPI_Window and MPI_Get
    time = Get<Test1, size>(world_size, world_rank, datatype2);
    if (world_rank == 0) std::cout << "{double, int, int} Get : " << time << " s" << std::endl;

    MPI_Finalize();
}

测试结果

mpirun -np 4 ./compareGetWithSendRecv 
{double, int} SendRecv : 0.0303547 s
{double, int} Get : 1.9196 s
{double, int, int} SendRecv : 0.0164659 s
{double, int, int} Get : 0.0147757 s

问题解答

1. 该结果是否正常?

是正常的,核心原因在于两种通信方式对非连续内存布局的处理逻辑不同:

  • MPI_SendRecv属于双边通信,发送端会根据自定义MPI类型,将结构体中的有效数据打包成连续的缓冲区后传输,接收端再解包还原,全程可以利用连续内存的批量传输优化(如DMA),因此效率不受结构体填充影响。
  • MPI_Get属于单边通信,直接从源进程的内存中读取数据。当结构体存在填充时,自定义MPI类型描述的是非连续的内存块(每个元素仅包含12字节有效数据,间隔4字节填充),MPI无法将多个元素合并为连续的内存区域读取,只能逐个处理非连续块,导致大量小数据传输,开销剧增。而无填充的Test2结构体内存布局连续,MPI_Get可以直接批量读取,性能与MPI_SendRecv接近。

2. 除添加冗余字段外的解决方案

(1)调整自定义MPI类型的Extent

通过MPI_Type_create_resized修改自定义类型的跨度(Extent),使其匹配结构体的实际内存大小(包含填充),让MPI识别出内存中连续的元素块,从而启用批量传输优化。

示例修改datatype1的创建逻辑:

MPI_Datatype datatype1_temp;
int array_of_blocklengths1[2] = {1, 1};
MPI_Aint array_of_displacements1[2] = {0, 8};
MPI_Datatype array_of_types1[2] = {MPI_DOUBLE, MPI_INT};
MPI_Type_create_struct(2, array_of_blocklengths1, array_of_displacements1, array_of_types1, &datatype1_temp);
// 将类型的Extent设置为结构体实际大小(16字节)
MPI_Type_create_resized(datatype1_temp, 0, sizeof(Test1), &datatype1);
MPI_Type_commit(&datatype1);
MPI_Type_free(&datatype1_temp);
(2)使用编译器指令控制结构体填充

通过编译器特性强制结构体无填充(如GCC的__attribute__((packed))、MSVC的#pragma pack(push, 1)),让结构体内存布局连续。但需注意:未对齐的内存访问在部分架构(如ARM)上会导致崩溃,在x86架构上也可能带来性能损耗,需权衡使用。

示例:

struct Test1 __attribute__((packed))
{
    double a;
    int b;
};
(3)转换为连续内存缓冲区传输

将结构体的有效数据复制到连续的数组中(如分离double和int到两个独立数组,或用字节数组打包),传输完成后再解析回结构体。这种方式完全规避了内存填充的影响,适用于所有场景。


内容的提问来源于stack exchange,提问作者Antoine Motte

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:17:00