相邻缓存行预取器缓存行数探究及实验偏差原因与改进方案
核心问题:相邻缓存行预取器会将多少条缓存行带入缓存?
我正在研究相邻缓存行预取器的有效性,以及它对从DRAM预取的缓存行数的影响。最初我假设它仅预取1条相邻缓存行。
我的目标是确定该预取器是否会预取缓存行,以及具体预取多少行。我仅在《Intel 64和IA-32架构软件开发手册》中找到关于0X1A4 MSR寄存器的相关信息。在我的Intel(R) Xeon(R) Gold 6442Y(第4代)处理器上,该寄存器值为0,表明所有预取器均已启用。
实验设计
- 分配一个1GB的uint64_t类型随机数组(超出我的128MB L3缓存),并用1-20的值初始化。
- 从初始化后的数组中读取8个随机索引,确保索引按4个缓存行对齐。
- 使用累加和取模操作随机选择已访问缓存行的相邻行。
- 引入变量以防止过度乱序执行。
- 使用GCC编译器的
-O2标志编译代码。
代码片段
#include <stdint.h> #include <stdio.h> #include <stdlib.h> #include <time.h> #define ARRAY_SIZE 135000000 // 1GB array #define CACHELINE_SIZE 64 #define ALIGN32(x) ((x) & ~(31)) #define ELEMENT_SIZE sizeof(uint64_t) int main(int argc, char *argv[]) { uint64_t *array = (uint64_t *)aligned_alloc(256, (ARRAY_SIZE * sizeof(uint64_t))); if (argc != 2) { printf("Usage: ./a.out <offset in Bytes> \n"); return 1; } int offset = atoi(argv[1]); int cacheline_offsets = (offset * ELEMENT_SIZE) / (CACHELINE_SIZE); printf("Number of cachelines offset is %d\n",cacheline_offsets); // Initialize the array for (int i = 0; i < ARRAY_SIZE; i++) { array[i] = rand() % 20; } // Measure execution time unsigned long sum = 0; uint64_t indices[8] = {0}; clock_t start = clock(); for (uint64_t i = 0; i < ARRAY_SIZE; i += 8) { uint64_t oldSum = sum; for (uint64_t j = 0; j < 8; j++) { indices[j] = (rand() ^ oldSum) % ARRAY_SIZE; indices[j] = ALIGN32(indices[j]); sum += array[indices[j]]; } sum += array[indices[sum % 8] + offset]; } clock_t end = clock(); double time_taken = ((double)(end - start)) / CLOCKS_PER_SEC; printf("Time taken: %f seconds\n", time_taken); printf("sum value: %ld\n", sum); // Use `sum` so it wont be optimized out free(array); return 0; }
实验结果

预期与实际行为对比
预期访问同一缓存行时性能最佳;预测访问硬件预取的相邻行(可能1-2条额外缓存行)时性能略有下降;预期访问更远的行时性能会进一步下降。但实际结果显示,后续7条相邻缓存行的访问性能一致,且性能下降幅度远低于预期(我原本认为DRAM miss会比L1 miss慢10倍)。
疑问
观察到的性能差异是否源于实验设计缺陷或错误假设?是否有其他方法可用于测量相邻缓存行预取器的影响?
内容的提问来源于stack exchange,提问作者Hod Badihi
相关产品推荐
相关产品推荐

