You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

循环体规模增大时IPC骤降但I-cache缺失率恒定的Skylake-SP CPU前端瓶颈排查问询

循环体规模增大时IPC骤降但I-cache缺失率恒定的Skylake-SP CPU前端瓶颈排查问询

我最近在做一组简单算术代码的性能测试,遇到了一个非常困惑的问题,想请教各位硬件性能优化的大佬:

当我逐步增大无分支循环体的代码规模时,每周期指令数(IPC)出现了大幅下滑,但L1指令缓存的缺失率却完全保持恒定。这显然不是指令缓存的问题,那会不会是CPU前端的其他瓶颈导致的?

现象细节

测试的是自动生成的纯算术代码,循环体里只有简单的64位加法操作,没有任何分支(除了循环本身的尾部分支,但其缺失率很低)。我用volatile修饰所有变量,确保编译器不会优化掉循环体里的操作。

当循环体行数从3万增加到10万时,IPC从2.08暴跌到1.30,但L1i缓存的缺失率始终稳定在9%左右,完全没有变化:

循环体行数IPCI-Cache缺失率
30K2.089.28%
40K1.949.27%
50K1.599.27%
60K1.449.27%
70K1.369.27%
100K1.309.27%

测试环境

  • CPU:Intel Xeon Gold 6132 @ 2.60GHz(Skylake-SP架构),双插槽,每插槽14核
    • L1i缓存:32KB/核
    • L2缓存:1MB/核
    • L3缓存:19.25MB(全插槽共享)
  • 编译器:gcc -O1
  • 测试工具:perf stat,每次测试用timeout 10s控制运行时长

补充的Perf性能数据

我额外采集了一些前端相关的性能计数器,其中:

  • idq_uops_not_delivered.cycles_fe_was_ok:1,489,045,825
  • resource_stalls.any:14,521,941
  • resource_stalls.sb:3,891,491

完整的70K行测试的perf输出:

perf stat -e cycles,instructions,branch-instructions,branch-misses,L1-dcache-load-misses,L1-icache-load-misses,LLC-load-misses,dTLB-load-misses,iTLB-load-misses -- timeout 10s ./test_70k
36,834,726,435 cycles
50,018,331,051 instructions # 1.36 insn per cycle
9,449,823 branch-instructions
321,923 branch-misses # 3.41% of all branches
327,729 L1-dcache-load-misses
4,638,663,649 L1-icache-load-misses
20,629 LLC-load-misses
15,639 dTLB-load-misses
154,250 iTLB-load-misses

测试代码说明

代码是自动生成的,核心结构如下:

#include <stdio.h>
#include <stdint.h>

// 用volatile防止编译器优化循环体操作
volatile uint64_t a = 1, b = 2, c = 3, d = 4;
volatile uint64_t e = 5, f = 6, g = 7, h = 8;
volatile uint64_t counter = 0;

int main() {
    printf("Starting 50000-line arithmetic test\n");
    while(1) {
        b = e + h;
        c = d + e;
        f = g + a;
        b = d + h;
        f = d + d;
        d = a + d;
        e = a + h;
        // ... 这里是数千行类似的64位加法操作
        c = e + b;
        h = a + b;
        e = c + e;
        f = b + d;
        // 这个分支几乎不会触发,只是为了避免循环被完全优化
        if ((a + b + c + d + e + f + g + h) == 0) {
            printf("Done\n");
        }
    }
    return 0;
}

代码生成逻辑很简单,用Python随机从8个volatile变量中选择操作数和结果变量,生成加法语句:

operations = ['+']
variables = ['a', 'b', 'c', 'd', 'e', 'f', 'g', 'h']
lines = []
for i in range(num_lines):
    var1 = random.choice(variables)
    var2 = random.choice(variables)
    result = random.choice(variables)
    op = random.choice(operations)
    lines.append(f"{result} = {var1} {op} {var2};")

核心疑问

  1. 既然I-cache的缺失率完全没有变化,排除了指令缓存的问题,那CPU前端的哪个瓶颈(解码带宽、微操作缓存uop cache、取指带宽)会导致这种IPC暴跌的现象?
  2. 我还需要采集哪些Perf性能计数器,才能精准定位到具体的瓶颈点?

麻烦各位大佬帮忙分析一下,感激不尽!

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.07 08:39:31