You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何矩阵乘法中输出数组c未出现预期的伪共享?

多线程矩阵乘法伪共享分析疑问解答

问题背景

在分析一个多线程矩阵乘法示例时,线程池中的不同线程会向输出数组c的不同4字节偏移位置写入计算结果(每个数组位置仅被写入一次)。根据伪共享的定义,预期会出现缓存行竞争,但perf c2c工具仅检测到原子计数器所在缓存行存在共享。

核心代码如下:

void fma_f32_single(const float* __restrict__ aptr,
                    const float* __restrict__ bptr, size_t M, size_t N,
                    size_t K, float* __restrict__ cptr) {
    float c0{0};
    for (size_t i = 0; i < K; ++i) {
        c0 += (*(aptr + i)) * (*(bptr + N * i));
    }
    *cptr = c0;
}

struct pthreadpool_context {
    const float* __restrict__ a;
    const float* __restrict__ b;
    float* __restrict__ c;
    size_t M;
    size_t N;
    size_t K;
    std::vector<std::pair<size_t, size_t>> indices;
};

void work(void* ctx, size_t i) {
    const pthreadpool_context* context = (pthreadpool_context*)ctx;
    const auto [row, col] = context->indices[i];
    const float* aptr = context->a + row * context->K;
    const float* bptr = context->b + col;
    float* cptr = context->c + row * context->N + col;
    // Increasing col by one is just a four byte different on c.
    // Threads write their output to cptr. 
    fma_f32_single(aptr, bptr, context->M, context->N, context->K, cptr);
}

使用perf c2c record -F 60000 ./a.out记录数据,perf c2c report -c tid,iaddr分析后的输出:

=================================================
            Trace Event Information              
=================================================
  Total records                     :     414185
  Locked Load/Store Operations      :        126
  Load Operations                   :     164311
  Loads - uncacheable               :          0
  Loads - IO                        :          0
  Loads - Miss                      :          1
  Loads - no mapping                :          7
  Load Fill Buffer Hit              :        477
  Load L1D hit                      :      84675
  Load L2D hit                      :          7
  Load LLC hit                      :      79115
  Load Local HITM                   :         40
  Load Remote HITM                  :          0
  Load Remote HIT                   :          0
  Load Local DRAM                   :         29
  Load Remote DRAM                  :          0
  Load MESI State Exclusive         :          0
  Load MESI State Shared            :         29
  Load LLC Misses                   :         29
  Load access blocked by data       :          0
  Load access blocked by address    :          0
  Load HIT Local Peer               :          0
  Load HIT Remote Peer              :          0
  LLC Misses to Local DRAM          :      100.0%
  LLC Misses to Remote DRAM         :        0.0%
  LLC Misses to Remote cache (HIT)  :        0.0%
  LLC Misses to Remote cache (HITM) :        0.0%
  Store Operations                  :     249874
  Store - uncacheable               :          0
  Store - no mapping                :          0
  Store L1D Hit                     :     249829
  Store L1D Miss                    :         45
  Store No available memory level   :          0
  No Page Map Rejects               :       4617
  Unable to parse data source       :          0

=================================================
    Global Shared Cache Line Event Information   
=================================================
  Total Shared Cache Lines          :          1
  Load HITs on shared lines         :        185
  Fill Buffer Hits on shared lines  :         43
  L1D hits on shared lines          :         66
  L2D hits on shared lines          :          0
  LLC hits on shared lines          :         76
  Load hits on peer cache or nodes  :          0
  Locked Access on shared lines     :        104
  Blocked Access on shared lines    :          0
  Store HITs on shared lines        :        785
  Store L1D hits on shared lines    :        785
  Store No available memory level   :          0
  Total Merged records              :        825

=================================================
                 c2c details                      
=================================================
  Events                            : cpu/mem-loads,ldlat=30/P
                                    : cpu/mem-stores/P
  Cachelines sort on                : Total HITMs
  Cacheline data grouping           : offset,tid,iaddr

=================================================
           Shared Data Cache Line Table          
=================================================
#
#        ----------- Cacheline ----------      Tot  ------- Load Hitm -------    Total    Total    Total  --------- Stores --------  ----- Core Load Hit -----  - LLC Load Hit --  - RMT Load Hit --  --- Load Dram ----
# Index             Address  Node  PA cnt     Hitm    Total  LclHitm  RmtHitm  records    Loads   Stores    L1Hit   L1Miss      N/A       FB       L1       L2    LclHit  LclHitm    RmtHit  RmtHitm       Lcl       Rmt
# .....  ..................  ....  ......  .......  .......  .......  .......  .......  .......  .......  .......  .......  .......  .......  .......  .......  ........  .......  ........  .......  ........  ........
#
      0      0x7fff33e60700     0     151  100.00%       40       40        0      970      185      785      785        0        0       43       66        0        36       40         0        0         0         0

=================================================
      Shared Cache Line Distribution Pareto      
=================================================
#
#        ----- HITM -----  ------- Store Refs ------  --------- Data address ---------                                     ---------- cycles ----------    Total       cpu                                  Shared                          
#   Num  RmtHitm  LclHitm   L1 Hit  L1 Miss      N/A              Offset  Node  PA cnt            Tid        Code address  rmt hitm  lcl hitm      load  records       cnt                          Symbol  Object        Source:Line  Node
# .....  .......  .......  .......  .......  .......  ..................  ....  ......  .............  ..................  ........  ........  ........  .......  ........  ..............................  .....  .................  ....
#
  ----------------------------------------------------------------------
      0        0       40      785        0        0      0x7fff33e60700
  ----------------------------------------------------------------------
           0.00%    2.50%   23.44%    0.00%    0.00%                0x34     0       1    84530:a.out      0x5e87bc063dfe         0       241       130      203         1  [.] ThreadPool::QueueTask(void  a.out  atomic_base.h:628   0
           0.00%    2.50%   24.84%    0.00%    0.00%                0x34     0       1    84533:a.out      0x5e87bc063a32         0       306       188      233         1  [.] ThreadMain(std::stop_token  a.out  atomic_base.h:628   0
           0.00%    0.00%   25.61%    0.00%    0.00%                0x34     0       1    84532:a.out      0x5e87bc063a32         0         0       167      225         1  [.] ThreadMain(std::stop_token  a.out  atomic_base.h:628   0
           0.00%    0.00%   26.11%    0.00%    0.00%                0x34     0       1    84534:a.out      0x5e87bc063a32         0         0       193      228         1  [.] ThreadMain(std::stop_token  a.out  atomic_base.h:628   0
           0.00%   42.50%    0.00%    0.00%    0.00%                0x38     0       1    84533:a.out      0x5e87bc0639f0         0       132       128       33         1  [.] ThreadMain(std::stop_token  a.out  threadpool.cpp:25   0
           0.00%   35.00%    0.00%    0.00%    0.00%                0x38     0       1    84534:a.out      0x5e87bc0639f0         0       133       121       28         1  [.] ThreadMain(std::stop_token  a.out  threadpool.cpp:25   0
           0.00%   17.50%    0.00%    0.00%    0.00%                0x38     0       1    84532:a.out      0x5e87bc0639f0         0       124       111       20         1  [.] ThreadMain(std::stop_token  a.out  threadpool.cpp:25   0

问题解答

1. 是否应该预期数组c出现伪共享?是否需要额外perf事件?

不应该预期出现伪共享,核心原因是伪共享的本质是多个线程对同一缓存行内的不同位置进行频繁的读写竞争,而你的场景中每个c数组元素仅被一个线程写入一次,且写入后没有任何线程再访问该元素:

  • 线程写入cptr时,会将对应缓存行加载到L1缓存并标记为Exclusive状态,写入后转为Modified状态;
  • 其他线程写入同一缓存行的其他元素时,会发起RFO(Read For Ownership)操作获取缓存行所有权,但由于之前的线程已经完成写入且不再访问该缓存行,不会出现来回的HITM(缓存行被修改后需要回写的情况)。perf c2c主要通过HITM事件识别伪共享,没有持续的HITM就不会标记该缓存行为共享。

另外,默认的perf c2c record已经覆盖了检测伪共享所需的mem-loads和mem-stores事件,无需额外添加事件。如果需要更细致的缓存行为分析,可以补充LLC-load-misses、LLC-store-misses等事件,但这不会改变当前场景的结果。

2. 其他架构(Intel其他型号、ARM)是否会出现此类伪共享?

伪共享是缓存行机制带来的通用问题,只要架构采用缓存行,存在持续的同一缓存行多线程读写竞争时就会出现伪共享,但你的场景在任何架构下都不会出现明显的伪共享问题,原因同上:每个元素仅被写入一次,没有后续的读写交互。

如果修改场景为多个线程反复读写同一缓存行的不同元素(比如循环更新c的元素),那么无论是Intel的Ice Lake、Sapphire Rapids,还是ARM的Neoverse N1/V1等架构,都会出现伪共享,对应的性能分析工具(如ARM上的perf mem)也能检测到相关的缓存行竞争。


内容的提问来源于stack exchange,提问作者fabian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 16:07:03