You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何带复杂寻址的mov指令比对应lea指令更快?

关于LEA与带复杂寻址MOV指令性能差异的困惑

我查阅了指令表,发现Coffee Lake架构中,含3个分量的lea指令RThroughput为1,我认为这速度很慢,因此推测带复杂寻址的mov指令RThroughput大于1。但令我惊讶的是,带复杂寻址的mov实际比lea更快,这让我十分困惑。

我的计算机微架构为Comet Lake,与Coffee Lake差异不大,以下是我使用的测试代码:

LEA指令测试代码(耗时8×10⁹ cycles)

mov ecx, 1000000000
xor rax, rax
sub rsp, 40
.align 32
loop:
    lea r8, [rsp + rax + 4]
    lea r9, [rsp + rax + 8]
    lea r10, [rsp + rax + 12]
    lea r11, [rsp + rax + 16]
    lea r12, [rsp + rax + 20]
    lea r13, [rsp + rax + 24]
    lea r14, [rsp + rax + 28]
    lea r15, [rsp + rax + 32]
    sub ecx, 1
    jnz loop
add rsp, 40

MOV指令测试代码(耗时4×10⁹ cycles)

mov ecx, 1000000000
xor rax, rax
sub rsp, 40
.align 32
loop:
    mov r8d, DWORD PTR [rsp + rax + 4]
    mov r9d, DWORD PTR [rsp + rax + 8]
    mov r10d, DWORD PTR [rsp + rax + 12]
    mov r11d, DWORD PTR [rsp + rax + 16]
    mov r12d, DWORD PTR [rsp + rax + 20]
    mov r13d, DWORD PTR [rsp + rax + 24]
    mov r14d, DWORD PTR [rsp + rax + 28]
    mov r15d, DWORD PTR [rsp + rax + 32]
    sub ecx, 1
    jnz loop
add rsp, 40

内容的提问来源于stack exchange,提问作者platelet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 22:37:25