为何矩阵乘法中输出数组c未出现预期的伪共享?
问题背景
在分析一个多线程矩阵乘法示例时,线程池中的不同线程会向输出数组c的不同4字节偏移位置写入计算结果(每个数组位置仅被写入一次)。根据伪共享的定义,预期会出现缓存行竞争,但perf c2c工具仅检测到原子计数器所在缓存行存在共享。
核心代码如下:
void fma_f32_single(const float* __restrict__ aptr, const float* __restrict__ bptr, size_t M, size_t N, size_t K, float* __restrict__ cptr) { float c0{0}; for (size_t i = 0; i < K; ++i) { c0 += (*(aptr + i)) * (*(bptr + N * i)); } *cptr = c0; } struct pthreadpool_context { const float* __restrict__ a; const float* __restrict__ b; float* __restrict__ c; size_t M; size_t N; size_t K; std::vector<std::pair<size_t, size_t>> indices; }; void work(void* ctx, size_t i) { const pthreadpool_context* context = (pthreadpool_context*)ctx; const auto [row, col] = context->indices[i]; const float* aptr = context->a + row * context->K; const float* bptr = context->b + col; float* cptr = context->c + row * context->N + col; // Increasing col by one is just a four byte different on c. // Threads write their output to cptr. fma_f32_single(aptr, bptr, context->M, context->N, context->K, cptr); }
使用perf c2c record -F 60000 ./a.out记录数据,perf c2c report -c tid,iaddr分析后的输出:
================================================= Trace Event Information ================================================= Total records : 414185 Locked Load/Store Operations : 126 Load Operations : 164311 Loads - uncacheable : 0 Loads - IO : 0 Loads - Miss : 1 Loads - no mapping : 7 Load Fill Buffer Hit : 477 Load L1D hit : 84675 Load L2D hit : 7 Load LLC hit : 79115 Load Local HITM : 40 Load Remote HITM : 0 Load Remote HIT : 0 Load Local DRAM : 29 Load Remote DRAM : 0 Load MESI State Exclusive : 0 Load MESI State Shared : 29 Load LLC Misses : 29 Load access blocked by data : 0 Load access blocked by address : 0 Load HIT Local Peer : 0 Load HIT Remote Peer : 0 LLC Misses to Local DRAM : 100.0% LLC Misses to Remote DRAM : 0.0% LLC Misses to Remote cache (HIT) : 0.0% LLC Misses to Remote cache (HITM) : 0.0% Store Operations : 249874 Store - uncacheable : 0 Store - no mapping : 0 Store L1D Hit : 249829 Store L1D Miss : 45 Store No available memory level : 0 No Page Map Rejects : 4617 Unable to parse data source : 0 ================================================= Global Shared Cache Line Event Information ================================================= Total Shared Cache Lines : 1 Load HITs on shared lines : 185 Fill Buffer Hits on shared lines : 43 L1D hits on shared lines : 66 L2D hits on shared lines : 0 LLC hits on shared lines : 76 Load hits on peer cache or nodes : 0 Locked Access on shared lines : 104 Blocked Access on shared lines : 0 Store HITs on shared lines : 785 Store L1D hits on shared lines : 785 Store No available memory level : 0 Total Merged records : 825 ================================================= c2c details ================================================= Events : cpu/mem-loads,ldlat=30/P : cpu/mem-stores/P Cachelines sort on : Total HITMs Cacheline data grouping : offset,tid,iaddr ================================================= Shared Data Cache Line Table ================================================= # # ----------- Cacheline ---------- Tot ------- Load Hitm ------- Total Total Total --------- Stores -------- ----- Core Load Hit ----- - LLC Load Hit -- - RMT Load Hit -- --- Load Dram ---- # Index Address Node PA cnt Hitm Total LclHitm RmtHitm records Loads Stores L1Hit L1Miss N/A FB L1 L2 LclHit LclHitm RmtHit RmtHitm Lcl Rmt # ..... .................. .... ...... ....... ....... ....... ....... ....... ....... ....... ....... ....... ....... ....... ....... ....... ........ ....... ........ ....... ........ ........ # 0 0x7fff33e60700 0 151 100.00% 40 40 0 970 185 785 785 0 0 43 66 0 36 40 0 0 0 0 ================================================= Shared Cache Line Distribution Pareto ================================================= # # ----- HITM ----- ------- Store Refs ------ --------- Data address --------- ---------- cycles ---------- Total cpu Shared # Num RmtHitm LclHitm L1 Hit L1 Miss N/A Offset Node PA cnt Tid Code address rmt hitm lcl hitm load records cnt Symbol Object Source:Line Node # ..... ....... ....... ....... ....... ....... .................. .... ...... ............. .................. ........ ........ ........ ....... ........ .............................. ..... ................. .... # ---------------------------------------------------------------------- 0 0 40 785 0 0 0x7fff33e60700 ---------------------------------------------------------------------- 0.00% 2.50% 23.44% 0.00% 0.00% 0x34 0 1 84530:a.out 0x5e87bc063dfe 0 241 130 203 1 [.] ThreadPool::QueueTask(void a.out atomic_base.h:628 0 0.00% 2.50% 24.84% 0.00% 0.00% 0x34 0 1 84533:a.out 0x5e87bc063a32 0 306 188 233 1 [.] ThreadMain(std::stop_token a.out atomic_base.h:628 0 0.00% 0.00% 25.61% 0.00% 0.00% 0x34 0 1 84532:a.out 0x5e87bc063a32 0 0 167 225 1 [.] ThreadMain(std::stop_token a.out atomic_base.h:628 0 0.00% 0.00% 26.11% 0.00% 0.00% 0x34 0 1 84534:a.out 0x5e87bc063a32 0 0 193 228 1 [.] ThreadMain(std::stop_token a.out atomic_base.h:628 0 0.00% 42.50% 0.00% 0.00% 0.00% 0x38 0 1 84533:a.out 0x5e87bc0639f0 0 132 128 33 1 [.] ThreadMain(std::stop_token a.out threadpool.cpp:25 0 0.00% 35.00% 0.00% 0.00% 0.00% 0x38 0 1 84534:a.out 0x5e87bc0639f0 0 133 121 28 1 [.] ThreadMain(std::stop_token a.out threadpool.cpp:25 0 0.00% 17.50% 0.00% 0.00% 0.00% 0x38 0 1 84532:a.out 0x5e87bc0639f0 0 124 111 20 1 [.] ThreadMain(std::stop_token a.out threadpool.cpp:25 0
问题解答
1. 是否应该预期数组c出现伪共享?是否需要额外perf事件?
不应该预期出现伪共享,核心原因是伪共享的本质是多个线程对同一缓存行内的不同位置进行频繁的读写竞争,而你的场景中每个c数组元素仅被一个线程写入一次,且写入后没有任何线程再访问该元素:
- 线程写入
cptr时,会将对应缓存行加载到L1缓存并标记为Exclusive状态,写入后转为Modified状态; - 其他线程写入同一缓存行的其他元素时,会发起RFO(Read For Ownership)操作获取缓存行所有权,但由于之前的线程已经完成写入且不再访问该缓存行,不会出现来回的HITM(缓存行被修改后需要回写的情况)。
perf c2c主要通过HITM事件识别伪共享,没有持续的HITM就不会标记该缓存行为共享。
另外,默认的perf c2c record已经覆盖了检测伪共享所需的mem-loads和mem-stores事件,无需额外添加事件。如果需要更细致的缓存行为分析,可以补充LLC-load-misses、LLC-store-misses等事件,但这不会改变当前场景的结果。
2. 其他架构(Intel其他型号、ARM)是否会出现此类伪共享?
伪共享是缓存行机制带来的通用问题,只要架构采用缓存行,存在持续的同一缓存行多线程读写竞争时就会出现伪共享,但你的场景在任何架构下都不会出现明显的伪共享问题,原因同上:每个元素仅被写入一次,没有后续的读写交互。
如果修改场景为多个线程反复读写同一缓存行的不同元素(比如循环更新c的元素),那么无论是Intel的Ice Lake、Sapphire Rapids,还是ARM的Neoverse N1/V1等架构,都会出现伪共享,对应的性能分析工具(如ARM上的perf mem)也能检测到相关的缓存行竞争。
内容的提问来源于stack exchange,提问作者fabian

