You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Intel CPU上的JMP指令会占用ALU端口?

无条件近跳转指令JMP的端口占用问题解析

我们来分析直接无条件近跳转指令JMP。它仅修改EIP寄存器,不消耗任何操作数(标志位、寄存器或内存),理论上完全可以在CPU前端完成解析。以下微基准测试显示,在Skylake架构CPU上,解析这类跳转耗时不到10个周期:

.globl _start
_start:
.rept 10000
        jmp .+3
        nop
.endr
        ud2
$ gcc -c t.s && ld t.o
$ perf stat -e cycles,instructions -e uops_dispatched_port.port_{0,1,5,6} ./a.out
./a.out: Illegal instruction

 Performance counter stats for './a.out':

            94,515      cycles:u
            10,002      instructions:u                   #    0.11  insn per cycle
               138       uops_dispatched_port.port_0:u
               152       uops_dispatched_port.port_1:u
               126       uops_dispatched_port.port_5:u
            10,232       uops_dispatched_port.port_6:u

       0.001461371 seconds time elapsed

       0.001111000 seconds user
       0.000000000 seconds sys

从测试结果可以看到,跳转指令会占用端口6,导致该端口无法被算术指令使用。这有点出乎意料,因为直接JMP似乎没有占用ALU端口的必要——毕竟Intel一直在致力于减少对执行端口的不必要占用:比如Sandy Bridge架构中,NOP和XOR same, same指令不占用任何端口;Ivy Bridge架构还新增了移动消除功能。

这个问题最终在Icelake架构上得到了修正。在该架构下,直接跳转指令不再占用端口,直接调用指令的执行微操作数也减少了一个。

那是否存在某些细微差异,使得JMP在处理器后端的处理逻辑和NOP不同,导致Intel直到Icelake才解决它占用ALU端口的问题?

内容的提问来源于stack exchange,提问作者amonakov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 00:38:24