You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

指针追踪基准测试:乱序执行未如预期生效问题排查

指针追踪基准测试(multichase)未按预期乱序执行的问题

我需要一个能产生大量缓存缺失的可靠基准测试,指针追踪测试是首选方案。尝试了google/multichase工具,但结果未达预期。

Multichase是经典指针追踪基准测试,作者做了额外处理确保硬件预取器失效,它会报告指针解引用的平均延迟,还支持插入空循环形式的"额外工作"。核心代码片段如下:

static void chase_work(per_thread_t *t) {
  void *p = t->x.cycle[0];
  size_t extra_work = strtoul(t->x.extra_args, 0, 0);
  size_t work = 0;
  size_t i;

  // the extra work is intended to be overlapped with a dereference,
  // but we don't want it to skip past the next dereference.  so
  // we fold in the value of the pointer, and launch the deref then
  // go into a loop performing extra work, hopefully while the
  // deref occurs.
  do {
    x25(work += (uintptr_t)p; p = *(void **)p;
        // (I uncomment the following line for the second benchmark)
        // (it was not a part of the original source code)
        //work += (uintptr_t)p;
        for (i = 0; i < extra_work; ++i) { work ^= i; })
  } while (__sync_add_and_fetch(&t->x.count, 25));

  // we never actually reach here, but the compiler doesn't know that
  t->x.cycle[0] = p;
  t->x.dummy = work;
}

(x25是重复参数25次的宏,即x25(y)等价于连续25次执行y)

extra_work为传入参数,理想情况下处理器应在等待指针取值时乱序执行额外工作循环。这意味着调整extra_work时,初期runtime应无明显增长,直到超过阈值后,runtime随extra_work线性增长(此时缓存缺失不再是瓶颈)。

但实际测试中并未出现初期的平缓阶段,全程呈线性增长,这让我怀疑循环并未在等待指针解引用时执行。

为验证这点,我在指针解引用后、循环前添加了一行work += (uintptr_t)p;,这会让解引用与循环强制有序,循环需等解引用完成才能执行,且原代码已确保下一次解引用需等循环结束。若此前循环确实与解引用乱序执行,修改后runtime应大幅上升,但实际仅增加了几纳秒。测试中extra_work取值1到500。

我的处理器为11代Intel Core i5-11400H,系统是Ubuntu 22.04。我认为循环理应与解引用乱序执行,但实际并未发生,因此有两个问题:

  1. 为何未出现预期的乱序执行?
  2. 如何修改才能让基准测试按预期工作?

复现步骤

  • 克隆multichase仓库
  • 编译后运行./multichase -c work:N(N为指定的extra_work数值)

内容的提问来源于stack exchange,提问作者Box Box Box Box

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 03:21:06