You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于perf stat报告分支数与实际代码分支数不符的原因问询

关于perf stat分支计数差异的疑问

我通过perf stat测试一个简单C++程序,代码如下:

#include <cassert>
#include <cstddef>
#include <iostream>

int main(int argc, const char* argv[]) {
    assert(argc == 3);
    int64_t iters = atoll(argv[1]);
    int64_t step = atoll(argv[2]);

    int64_t value = 0;

    for(int64_t i = 0; i < iters; ++i) {
        value += step;
    }
    std::cout << value << std::endl;

    return 0;
}

从汇编代码中可以看到,循环仅包含一个重复分支指令jge .LBB0_4。

使用的编译命令为:

clang++ -std=c++2b bp.cpp tp2.cpp -o bp.exe -Wall -O0 -DNDEBUG

执行以下perf stat命令:

perf stat ./bp.exe 10000000 700 500 1013 2>&1 | grep branches | tee run4_1.txt
perf stat ./bp.exe 20000000 700 500 1013 2>&1 | grep branches | tee run4_2.txt
perf stat ./bp.exe 30000000 700 500 1013 2>&1 | grep branches | tee run4_3.txt

输出结果为:

+ perf stat ./bp_arc.exe 10000000 700
+ grep branches
+ tee run4_1_arc.txt
          15516156      branches:u                #  483.498 M/sec                    (66.64%)
              4390      branch-misses:u           #    0.03% of all branches          (63.53%)
+ perf stat ./bp_arc.exe 20000000 700
+ grep branches
+ tee run4_2_arc.txt
          30860233      branches:u                #  534.704 M/sec                    (67.69%)
              6400      branch-misses:u           #    0.02% of all branches          (65.99%)
+ perf stat ./bp_arc.exe 30000000 700
+ grep branches
+ tee run4_3_arc.txt
          47616449      branches:u                #  535.042 M/sec                    (67.21%)
              6531      branch-misses:u           #    0.01% of all branches          (66.81%)

可见perf stat报告的分支数约为迭代次数的1.5倍,在其他环境下该倍数约为2.1倍。现咨询:perf stat的branches计数器统计的是哪些分支,为何会与汇编中可见的分支检查数存在差异?


内容的提问来源于stack exchange,提问作者ilnurKh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.01 18:37:29