You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何更多x86指令的执行速度反而快于更少指令?——C编译代码与手动汇编代码的性能差异问询

Why Your 7-Instruction Assembly Loop Is Slower Than the Compiler's 18-Instruction Version

Great question—this is exactly the kind of deep dive that teaches you how modern x86 CPUs really work, beyond just counting instruction counts. Let’s break down the key reasons your hand-written assembly is underperforming:

1. Your Code Has a Long Serial Dependency Chain (No Parallelism)

The biggest issue is that you’re over-reusing the ax register, creating a rigid chain where every instruction depends on the result of the previous one. Let’s walk through your loop’s core logic:

movzx ax,byte ptr [esi]       ; ax = source pixel
mov ah,byte ptr renderer_light[...]  ; ah = lightLUT[ax][light] (overwrites ah)
mov al,byte ptr [edi]         ; al = dest pixel (overwrites al)
mov al,byte ptr renderer_trans[...]  ; al = transLUT[ah][al][trans] (overwrites al again)
mov byte ptr [edi],al         ; write result to dest

Every step here waits for the prior one to finish. Modern x86 CPUs are 超标量(superscalar) and 乱序执行(out-of-order)—they can run multiple instructions at the same time, but only if those instructions don’t depend on each other’s results. Your code gives the CPU zero room to parallelize work.

Compare this to the compiler’s output: it uses eax, ecx, and edx to split up intermediate values. For example, while it’s waiting for the renderer_light lookup to complete, it can already start reading the destination pixel and calculating the renderer_trans address. These independent operations run in parallel, hiding memory latency and using the CPU’s execution units more efficiently.

2. The loop Instruction Is Slow on Modern CPUs

Your loop drawing_loop might seem efficient (one instruction instead of dec ecx / jnz), but it’s actually a legacy instruction that performs poorly on modern x86 microarchitectures. Here’s why:

  • loop combines two operations (decrement ecx and check for zero) into a single instruction, but CPUs can’t split this into separate micro-ops to execute in parallel with other work.
  • It also has higher latency than the manual dec ecx / jnz pair, and branch predictors often handle simple jnz branches better than loop.

The compiler avoids loop entirely for good reason—its "single instruction" convenience comes at the cost of worse performance.

3. 16-Bit Register Operations Have Hidden Overhead

You’re using the 16-bit ax register, but in 32-bit mode, 16-bit register accesses can trigger extra work for the CPU’s register renaming logic. When you modify ah or al, the CPU has to track the upper/lower bits of eax separately, which can introduce small but cumulative delays. The compiler’s code uses 32-bit registers exclusively, avoiding this overhead.

4. Memory Prefetching Can’t Help Your Code

CPUs use prefetchers to guess which memory addresses you’ll need next and load them into cache ahead of time. But prefetchers rely on predictable, independent memory access patterns.

In your code, the next LUT lookup address depends entirely on the result of the previous instruction (e.g., the renderer_trans address needs the ah value from the renderer_light lookup). The prefetcher can’t guess what that address will be until the prior instruction finishes, so it can’t hide the memory access latency.

The compiler’s code, by contrast, has independent memory operations. While it’s waiting for one LUT lookup to complete, it can already start prefetching the next one, or load the destination pixel from memory—turning idle waiting time into useful work.

Quick Fixes to Speed Up Your Assembly

Try these tweaks to match (or beat) the compiler’s performance:

  • Split up register usage: Use separate 32-bit registers for each intermediate value (e.g., ecx for source pixel, edx for light-adjusted color, eax for destination pixel). Break the dependency chain so the CPU can parallelize work.
  • Replace loop with dec ecx / jnz: Swap loop drawing_loop for:
    dec ecx
    jnz drawing_loop
    
  • Use 32-bit registers exclusively: Ditch ax and use eax/ecx/edx for all operations to avoid 16-bit overhead.
  • Rearrange instructions: Move independent operations (like inc esi/inc edi) earlier in the loop if possible, so they can run in parallel with memory lookups.

内容的提问来源于stack exchange,提问作者Napoleon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 13:54:11