You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Xeon平台随机访存代码L1 d-cache失效率偏低原因探究

L1 d-cache失效率测试异常偏低问题分析

我最初尝试编写刻意构造极低L1 d-cache命中率的测试代码,初始版本如下:

#include <stdio.h>
#include <sys/time.h>
#include <stdlib.h>

#define S 16*1024*1024
int largedata[S];

int main()
{
    struct timeval tv1;
    struct timeval tv2;

    for (int j = 0; j < 10; j++)
    {
            gettimeofday(&tv1, NULL);
            int total = 0;
            for (int i = 1; i < S; i++) {
                    largedata[i] = largedata[rand()%S] + 1;
                    total += largedata[i];
            }
            gettimeofday(&tv2, NULL);
            int elapsed = 1000000 * (tv2.tv_sec - tv1.tv_sec) + (tv2.tv_usec - tv1.tv_usec);
            printf("Round %d elapsed %d us --> %d\n", j, elapsed, total);
    }

    return 0;
}

这段代码定义了大小为64MB的缓冲区,对其执行随机访问。测试所用机器的L1 D-cache大小为32KB,原本预期会得到极低的缓存命中率(即20%~30%甚至更高的失效率),但实测得到的L1 d-cache失效率仅为2.57%,perf统计结果如下:

# gcc test.c &&    perf stat -e L1-dcache-load-misses -e L1-icache-load-misses -e L1-dcache-loads -e L1-dcache-stores  ./a.out
Round 0 elapsed 399706 us --> 26671607
Round 1 elapsed 344664 us --> 57210118
Round 2 elapsed 342444 us --> 79296375
Round 3 elapsed 344605 us --> 92293029
Round 4 elapsed 342173 us --> 93904234
Round 5 elapsed 346295 us --> 76478386
Round 6 elapsed 343390 us --> 98878844
Round 7 elapsed 347893 us --> 107286968
Round 8 elapsed 355442 us --> 87289283
Round 9 elapsed 362253 us --> 101320374

 Performance counter stats for './a.out':

     104128257      L1-dcache-load-misses     #    2.57% of all L1-dcache accesses
       2630683      L1-icache-load-misses
    4047961891      L1-dcache-loads
    2192632892      L1-dcache-stores

   3.539076619 seconds time elapsed

   3.479520000 seconds user
   0.032746000 seconds sys

测试使用的CPU为Intel(R) Xeon(R) Gold 6230,需要明确x86架构Xeon处理器出现该现象的原因。


问题根因

失效率数据异常偏低的核心原因是rand()函数的开销被统计进了总L1缓存访问计数,稀释了随机访问大数组产生的缓存失效占比:

  • 标准库rand()函数内部维护状态、执行计算的过程中,每次调用会产生约30次内存访问,这些访问几乎全部命中L1缓存,导致总L1 d-cache访问量被大幅拉高
  • 测试关注的大数组随机访问产生的缓存失效,在总访问量中的占比被这些高频命中的访问拉低,最终呈现出仅2.57%的失效率假象,和CPU硬件缓存机制无关

验证与修正

将随机数生成逻辑替换为无额外内存访问开销的xorshift轻量随机算法后,核心循环代码如下:

for (int j = 0; j < 10; j++)
    {
            gettimeofday(&tv1, NULL);
            int total = 0;
            for (int i = 1; i < S; i++) {

                    uint32_t t = x;
                    t ^= t << 11U;
                    t ^= t >> 8U;
                    x = y; y = z; z = w;
                    w ^= w >> 19U;
                    w ^= t;

                    largedata[i] = largedata[w%S] + 1;
                    total += largedata[i];
            }
            gettimeofday(&tv2, NULL);
            int elapsed = 1000000 * (tv2.tv_sec - tv1.tv_sec) + (tv2.tv_usec - tv1.tv_usec);
            printf("Round %d elapsed %d us --> %d\n", j, elapsed, total);
    }

修正后实测L1 d-cache失效率达到47%,符合随机访问远超L1容量内存的预期结果,perf统计数据如下:

Performance counter stats for './a.out':

      87715381      L1-dcache-load-misses     #   47.35% of all L1-dcache accesses
       1064131      L1-icache-load-misses
     185235862      L1-dcache-loads
     177250983      L1-dcache-stores

   0.715463900 seconds time elapsed

   0.664886000 seconds user
   0.028505000 seconds sys

对比两组数据可以看到,替换随机算法后总L1-dcache-loads从40亿次降到了1.85亿次,减少的部分几乎全是rand()调用产生的冗余L1命中访问,直接验证了之前失效率被稀释的判断。


内容的提问来源于stack exchange,提问作者jaeyong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 06:21:30