You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

相邻缓存行预取器缓存行数探究及实验偏差原因与改进方案

核心问题:相邻缓存行预取器会将多少条缓存行带入缓存?

我正在研究相邻缓存行预取器的有效性,以及它对从DRAM预取的缓存行数的影响。最初我假设它仅预取1条相邻缓存行。

我的目标是确定该预取器是否会预取缓存行,以及具体预取多少行。我仅在《Intel 64和IA-32架构软件开发手册》中找到关于0X1A4 MSR寄存器的相关信息。在我的Intel(R) Xeon(R) Gold 6442Y(第4代)处理器上,该寄存器值为0,表明所有预取器均已启用。

实验设计

  • 分配一个1GB的uint64_t类型随机数组(超出我的128MB L3缓存),并用1-20的值初始化。
  • 从初始化后的数组中读取8个随机索引,确保索引按4个缓存行对齐。
  • 使用累加和取模操作随机选择已访问缓存行的相邻行。
  • 引入变量以防止过度乱序执行。
  • 使用GCC编译器的-O2标志编译代码。

代码片段

#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <time.h>

#define ARRAY_SIZE 135000000 // 1GB array
#define CACHELINE_SIZE 64
#define ALIGN32(x) ((x) & ~(31))
#define ELEMENT_SIZE sizeof(uint64_t)

int main(int argc, char *argv[]) {
  uint64_t *array =
      (uint64_t *)aligned_alloc(256, (ARRAY_SIZE * sizeof(uint64_t)));

  if (argc != 2) {
    printf("Usage: ./a.out <offset in Bytes> \n");
    return 1;
  }

  int offset = atoi(argv[1]);
  int cacheline_offsets = (offset * ELEMENT_SIZE) / (CACHELINE_SIZE);
  printf("Number of cachelines offset is %d\n",cacheline_offsets);

  // Initialize the array
  for (int i = 0; i < ARRAY_SIZE; i++) {
    array[i] = rand() % 20;
  }

  // Measure execution time
  unsigned long sum = 0;
  uint64_t indices[8] = {0};

  clock_t start = clock();
  for (uint64_t i = 0; i < ARRAY_SIZE; i += 8) {
    uint64_t oldSum = sum;
    for (uint64_t j = 0; j < 8; j++) {
      indices[j] = (rand() ^ oldSum) % ARRAY_SIZE;
      indices[j] = ALIGN32(indices[j]);
      sum += array[indices[j]];
    }
    sum += array[indices[sum % 8] + offset];
  }
  clock_t end = clock();

  double time_taken = ((double)(end - start)) / CLOCKS_PER_SEC;
  printf("Time taken: %f seconds\n", time_taken);
  printf("sum value: %ld\n", sum); // Use `sum` so it wont be optimized out

  free(array);
  return 0;
}

实验结果

实验结果图

预期与实际行为对比

预期访问同一缓存行时性能最佳;预测访问硬件预取的相邻行(可能1-2条额外缓存行)时性能略有下降;预期访问更远的行时性能会进一步下降。但实际结果显示,后续7条相邻缓存行的访问性能一致,且性能下降幅度远低于预期(我原本认为DRAM miss会比L1 miss慢10倍)。

疑问

观察到的性能差异是否源于实验设计缺陷或错误假设?是否有其他方法可用于测量相邻缓存行预取器的影响?


内容的提问来源于stack exchange,提问作者Hod Badihi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 14:04:56