指针追踪基准测试:乱序执行未如预期生效问题排查
指针追踪基准测试(multichase)未按预期乱序执行的问题
我需要一个能产生大量缓存缺失的可靠基准测试,指针追踪测试是首选方案。尝试了google/multichase工具,但结果未达预期。
Multichase是经典指针追踪基准测试,作者做了额外处理确保硬件预取器失效,它会报告指针解引用的平均延迟,还支持插入空循环形式的"额外工作"。核心代码片段如下:
static void chase_work(per_thread_t *t) { void *p = t->x.cycle[0]; size_t extra_work = strtoul(t->x.extra_args, 0, 0); size_t work = 0; size_t i; // the extra work is intended to be overlapped with a dereference, // but we don't want it to skip past the next dereference. so // we fold in the value of the pointer, and launch the deref then // go into a loop performing extra work, hopefully while the // deref occurs. do { x25(work += (uintptr_t)p; p = *(void **)p; // (I uncomment the following line for the second benchmark) // (it was not a part of the original source code) //work += (uintptr_t)p; for (i = 0; i < extra_work; ++i) { work ^= i; }) } while (__sync_add_and_fetch(&t->x.count, 25)); // we never actually reach here, but the compiler doesn't know that t->x.cycle[0] = p; t->x.dummy = work; }
(x25是重复参数25次的宏,即x25(y)等价于连续25次执行y)
extra_work为传入参数,理想情况下处理器应在等待指针取值时乱序执行额外工作循环。这意味着调整extra_work时,初期runtime应无明显增长,直到超过阈值后,runtime随extra_work线性增长(此时缓存缺失不再是瓶颈)。
但实际测试中并未出现初期的平缓阶段,全程呈线性增长,这让我怀疑循环并未在等待指针解引用时执行。
为验证这点,我在指针解引用后、循环前添加了一行work += (uintptr_t)p;,这会让解引用与循环强制有序,循环需等解引用完成才能执行,且原代码已确保下一次解引用需等循环结束。若此前循环确实与解引用乱序执行,修改后runtime应大幅上升,但实际仅增加了几纳秒。测试中extra_work取值1到500。
我的处理器为11代Intel Core i5-11400H,系统是Ubuntu 22.04。我认为循环理应与解引用乱序执行,但实际并未发生,因此有两个问题:
- 为何未出现预期的乱序执行?
- 如何修改才能让基准测试按预期工作?
复现步骤
- 克隆multichase仓库
- 编译后运行
./multichase -c work:N(N为指定的extra_work数值)
内容的提问来源于stack exchange,提问作者Box Box Box Box
相关产品推荐
相关产品推荐

