循环体规模增大时IPC骤降但I-cache缺失率恒定的Skylake-SP CPU前端瓶颈排查问询
循环体规模增大时IPC骤降但I-cache缺失率恒定的Skylake-SP CPU前端瓶颈排查问询
我最近在做一组简单算术代码的性能测试,遇到了一个非常困惑的问题,想请教各位硬件性能优化的大佬:
当我逐步增大无分支循环体的代码规模时,每周期指令数(IPC)出现了大幅下滑,但L1指令缓存的缺失率却完全保持恒定。这显然不是指令缓存的问题,那会不会是CPU前端的其他瓶颈导致的?
现象细节
测试的是自动生成的纯算术代码,循环体里只有简单的64位加法操作,没有任何分支(除了循环本身的尾部分支,但其缺失率很低)。我用volatile修饰所有变量,确保编译器不会优化掉循环体里的操作。
当循环体行数从3万增加到10万时,IPC从2.08暴跌到1.30,但L1i缓存的缺失率始终稳定在9%左右,完全没有变化:
| 循环体行数 | IPC | I-Cache缺失率 |
|---|---|---|
| 30K | 2.08 | 9.28% |
| 40K | 1.94 | 9.27% |
| 50K | 1.59 | 9.27% |
| 60K | 1.44 | 9.27% |
| 70K | 1.36 | 9.27% |
| 100K | 1.30 | 9.27% |
测试环境
- CPU:Intel Xeon Gold 6132 @ 2.60GHz(Skylake-SP架构),双插槽,每插槽14核
- L1i缓存:32KB/核
- L2缓存:1MB/核
- L3缓存:19.25MB(全插槽共享)
- 编译器:gcc -O1
- 测试工具:
perf stat,每次测试用timeout 10s控制运行时长
补充的Perf性能数据
我额外采集了一些前端相关的性能计数器,其中:
idq_uops_not_delivered.cycles_fe_was_ok:1,489,045,825resource_stalls.any:14,521,941resource_stalls.sb:3,891,491
完整的70K行测试的perf输出:
perf stat -e cycles,instructions,branch-instructions,branch-misses,L1-dcache-load-misses,L1-icache-load-misses,LLC-load-misses,dTLB-load-misses,iTLB-load-misses -- timeout 10s ./test_70k 36,834,726,435 cycles 50,018,331,051 instructions # 1.36 insn per cycle 9,449,823 branch-instructions 321,923 branch-misses # 3.41% of all branches 327,729 L1-dcache-load-misses 4,638,663,649 L1-icache-load-misses 20,629 LLC-load-misses 15,639 dTLB-load-misses 154,250 iTLB-load-misses
测试代码说明
代码是自动生成的,核心结构如下:
#include <stdio.h> #include <stdint.h> // 用volatile防止编译器优化循环体操作 volatile uint64_t a = 1, b = 2, c = 3, d = 4; volatile uint64_t e = 5, f = 6, g = 7, h = 8; volatile uint64_t counter = 0; int main() { printf("Starting 50000-line arithmetic test\n"); while(1) { b = e + h; c = d + e; f = g + a; b = d + h; f = d + d; d = a + d; e = a + h; // ... 这里是数千行类似的64位加法操作 c = e + b; h = a + b; e = c + e; f = b + d; // 这个分支几乎不会触发,只是为了避免循环被完全优化 if ((a + b + c + d + e + f + g + h) == 0) { printf("Done\n"); } } return 0; }
代码生成逻辑很简单,用Python随机从8个volatile变量中选择操作数和结果变量,生成加法语句:
operations = ['+'] variables = ['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h'] lines = [] for i in range(num_lines): var1 = random.choice(variables) var2 = random.choice(variables) result = random.choice(variables) op = random.choice(operations) lines.append(f"{result} = {var1} {op} {var2};")
核心疑问
- 既然I-cache的缺失率完全没有变化,排除了指令缓存的问题,那CPU前端的哪个瓶颈(解码带宽、微操作缓存uop cache、取指带宽)会导致这种IPC暴跌的现象?
- 我还需要采集哪些Perf性能计数器,才能精准定位到具体的瓶颈点?
麻烦各位大佬帮忙分析一下,感激不尽!
内容来源于stack exchange
相关产品推荐
相关产品推荐

