关于perf stat报告分支数与实际代码分支数不符的原因问询
关于perf stat分支计数差异的疑问
我通过perf stat测试一个简单C++程序,代码如下:
#include <cassert> #include <cstddef> #include <iostream> int main(int argc, const char* argv[]) { assert(argc == 3); int64_t iters = atoll(argv[1]); int64_t step = atoll(argv[2]); int64_t value = 0; for(int64_t i = 0; i < iters; ++i) { value += step; } std::cout << value << std::endl; return 0; }
从汇编代码中可以看到,循环仅包含一个重复分支指令jge .LBB0_4。
使用的编译命令为:
clang++ -std=c++2b bp.cpp tp2.cpp -o bp.exe -Wall -O0 -DNDEBUG
执行以下perf stat命令:
perf stat ./bp.exe 10000000 700 500 1013 2>&1 | grep branches | tee run4_1.txt perf stat ./bp.exe 20000000 700 500 1013 2>&1 | grep branches | tee run4_2.txt perf stat ./bp.exe 30000000 700 500 1013 2>&1 | grep branches | tee run4_3.txt
输出结果为:
+ perf stat ./bp_arc.exe 10000000 700 + grep branches + tee run4_1_arc.txt 15516156 branches:u # 483.498 M/sec (66.64%) 4390 branch-misses:u # 0.03% of all branches (63.53%) + perf stat ./bp_arc.exe 20000000 700 + grep branches + tee run4_2_arc.txt 30860233 branches:u # 534.704 M/sec (67.69%) 6400 branch-misses:u # 0.02% of all branches (65.99%) + perf stat ./bp_arc.exe 30000000 700 + grep branches + tee run4_3_arc.txt 47616449 branches:u # 535.042 M/sec (67.21%) 6531 branch-misses:u # 0.01% of all branches (66.81%)
可见perf stat报告的分支数约为迭代次数的1.5倍,在其他环境下该倍数约为2.1倍。现咨询:perf stat的branches计数器统计的是哪些分支,为何会与汇编中可见的分支检查数存在差异?
内容的提问来源于stack exchange,提问作者ilnurKh
相关产品推荐
相关产品推荐

