You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何遍历单个结构体全数据时SOA仍比AOS更快?

SOA与AOS性能测试结果与预期不符的原因分析

问题描述

我编写了如下C++代码测试SOA(Struct of Array)与AOS(Array of Struct)的性能:

#include <chrono>
#include <array>
#include <fstream>
#include <iostream>
#include <istream>
#include <ostream>
#include <random>
#include <stdio.h>
#include <vector>

#define ELE_NUM 50000000

struct SOA{
    std::vector<int> array1;
    std::vector<int> array2;
    std::vector<int> array3;
};

struct AOS{
    int a;
    int b;
    int c;
};


int main(){
    
    std::random_device _rd;
    std::mt19937 rd(_rd());
    std::uniform_int_distribution<int> ds(0,100000);
    SOA soa;
    std::vector<AOS> aos(ELE_NUM);
    std::vector<int> rst(ELE_NUM);
    soa.array1.reserve(ELE_NUM);
    soa.array2.reserve(ELE_NUM);
    soa.array3.reserve(ELE_NUM); 
    auto begin_time = std::chrono::system_clock::now();
    for(int i = 0 ; i < ELE_NUM;++i){
        aos[i].a = ds(rd);
        aos[i].b = ds(rd);
        aos[i].c = ds(rd);
    }
    auto end_time = std::chrono::system_clock::now();
    auto aos_write_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time-begin_time);

    begin_time = std::chrono::system_clock::now();
    for(int i = 0 ; i < ELE_NUM;++i){
        soa.array1[i] = ds(rd);
        soa.array2[i] = ds(rd);
        soa.array3[i] = ds(rd);
    }
    end_time = std::chrono::system_clock::now();
    auto soa_write_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time-begin_time);

    begin_time = std::chrono::system_clock::now();
    for(int i = 0 ; i < ELE_NUM ; i++){
        rst[i] = aos[i].a +aos[i].b +aos[i].c;
    }
    end_time = std::chrono::system_clock::now();
    auto aos_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time-begin_time);
    
    //for avoid optimize, output rst
    std::fstream f("tmp.txt",std::iostream::out);
    for(int i = 0 ; i < ELE_NUM ; ++i ) f<<rst[i];
    f<<std::endl;
    begin_time = std::chrono::system_clock::now();

    for (int i = 0 ; i < ELE_NUM; i++) {
        rst[i] = soa.array1[i] + soa.array2[i] + soa.array3[i];
    }
    end_time = std::chrono::system_clock::now();
    auto soa_time = std::chrono::duration_cast<std::chrono::milliseconds>(end_time-begin_time);

    for(int i = 0 ; i < ELE_NUM ; ++i ) f<<rst[i];
    f<<std::endl;
    f.close();

    std::cout<<"soa_write_time:"<<soa_write_time.count()<<" ms"<<std::endl;
    std::cout<<"aos_write_time:"<<aos_write_time.count()<<" ms"<<std::endl;
    std::cout<<"soa_time:"<<soa_time.count()<<" ms"<<std::endl;
    std::cout<<"aos_time:"<<aos_time.count()<<" ms"<<std::endl;

    return 0;
}

测试环境:Windows系统、i5 13400处理器,使用g++ -O3编译
输出结果:

soa_write_time:225 ms
aos_write_time:179 ms
soa_time:28 ms
aos_time:43 ms

按我的理解,SOA将单个结构体的数据存储在独立数组中,遍历单个结构体全数据时无法命中CPU缓存,应比AOS慢;而AOS将结构体数据连续存储,能命中CPU缓存,应更快,但测试结果相反,请问这是为什么?


原因分析

1. 读取阶段:SOA适配批量字段操作的缓存效率更高

你的读取逻辑是对所有元素的三个字段分别求和,这种访问模式下:

  • SOA的array1、array2、array3都是连续内存块。遍历每个数组时,CPU预取机制会加载连续的缓存行,整个访问过程缓存命中率极高,几乎不需要从内存重新加载数据。
  • AOS的存储是单个结构体的三个字段连续,但读取时需要跨结构体访问同名字段(比如所有a字段),属于跨步访问(每次跳过12字节)。此时CPU预取的缓存行中,大部分是当前不需要的b、c字段,缓存利用率极低,大量数据需要从内存读取,导致速度变慢。

2. 写入阶段:AOS的连续内存访问更高效

写入逻辑是给单个元素的三个字段依次赋值,这种场景下:

  • AOS的整个vector<AOS>是一块连续内存,写入时是连续地址访问,CPU可以高效预取和批量写入,内存总线利用率最大化。
  • SOA需要在三个独立的vector之间切换写入,频繁的内存地址跳转破坏了访问连续性,导致CPU预取机制失效,同时增加了内存控制器的调度开销,最终写入速度慢于AOS。

3. SOA代码存在未定义行为

你的SOA仅调用了reserve(ELE_NUM),但未初始化元素。vector::reserve只预留内存空间,不会构造元素,直接访问soa.array1[i]属于未定义行为。虽然测试中未崩溃,但这种操作可能引入额外的内存访问开销,影响SOA写入性能的准确性。正确做法是调用resize(ELE_NUM)初始化元素。

核心结论

SOA和AOS的性能优劣完全取决于访问模式:

  • 当需要对同一字段做批量操作时,SOA的连续内存访问能最大化缓存效率,性能更优;
  • 当需要对单个结构体的所有字段做操作时,AOS的连续存储能提升缓存命中率,性能更好。

你的测试场景正好对应这两种情况:读取是批量字段操作,SOA占优;写入是单结构体多字段赋值,AOS占优,因此结果符合两者的特性,和预期不符是因为混淆了场景对应的访问模式。


内容的提问来源于stack exchange,提问作者Dreams

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 17:35:55