为何更多x86指令的执行速度反而快于更少指令?——C编译代码与手动汇编代码的性能差异问询
Great question—this is exactly the kind of deep dive that teaches you how modern x86 CPUs really work, beyond just counting instruction counts. Let’s break down the key reasons your hand-written assembly is underperforming:
1. Your Code Has a Long Serial Dependency Chain (No Parallelism)
The biggest issue is that you’re over-reusing the ax register, creating a rigid chain where every instruction depends on the result of the previous one. Let’s walk through your loop’s core logic:
movzx ax,byte ptr [esi] ; ax = source pixel mov ah,byte ptr renderer_light[...] ; ah = lightLUT[ax][light] (overwrites ah) mov al,byte ptr [edi] ; al = dest pixel (overwrites al) mov al,byte ptr renderer_trans[...] ; al = transLUT[ah][al][trans] (overwrites al again) mov byte ptr [edi],al ; write result to dest
Every step here waits for the prior one to finish. Modern x86 CPUs are 超标量(superscalar) and 乱序执行(out-of-order)—they can run multiple instructions at the same time, but only if those instructions don’t depend on each other’s results. Your code gives the CPU zero room to parallelize work.
Compare this to the compiler’s output: it uses eax, ecx, and edx to split up intermediate values. For example, while it’s waiting for the renderer_light lookup to complete, it can already start reading the destination pixel and calculating the renderer_trans address. These independent operations run in parallel, hiding memory latency and using the CPU’s execution units more efficiently.
2. The loop Instruction Is Slow on Modern CPUs
Your loop drawing_loop might seem efficient (one instruction instead of dec ecx / jnz), but it’s actually a legacy instruction that performs poorly on modern x86 microarchitectures. Here’s why:
loopcombines two operations (decrementecxand check for zero) into a single instruction, but CPUs can’t split this into separate micro-ops to execute in parallel with other work.- It also has higher latency than the manual
dec ecx / jnzpair, and branch predictors often handle simplejnzbranches better thanloop.
The compiler avoids loop entirely for good reason—its "single instruction" convenience comes at the cost of worse performance.
3. 16-Bit Register Operations Have Hidden Overhead
You’re using the 16-bit ax register, but in 32-bit mode, 16-bit register accesses can trigger extra work for the CPU’s register renaming logic. When you modify ah or al, the CPU has to track the upper/lower bits of eax separately, which can introduce small but cumulative delays. The compiler’s code uses 32-bit registers exclusively, avoiding this overhead.
4. Memory Prefetching Can’t Help Your Code
CPUs use prefetchers to guess which memory addresses you’ll need next and load them into cache ahead of time. But prefetchers rely on predictable, independent memory access patterns.
In your code, the next LUT lookup address depends entirely on the result of the previous instruction (e.g., the renderer_trans address needs the ah value from the renderer_light lookup). The prefetcher can’t guess what that address will be until the prior instruction finishes, so it can’t hide the memory access latency.
The compiler’s code, by contrast, has independent memory operations. While it’s waiting for one LUT lookup to complete, it can already start prefetching the next one, or load the destination pixel from memory—turning idle waiting time into useful work.
Quick Fixes to Speed Up Your Assembly
Try these tweaks to match (or beat) the compiler’s performance:
- Split up register usage: Use separate 32-bit registers for each intermediate value (e.g.,
ecxfor source pixel,edxfor light-adjusted color,eaxfor destination pixel). Break the dependency chain so the CPU can parallelize work. - Replace
loopwithdec ecx / jnz: Swaploop drawing_loopfor:dec ecx jnz drawing_loop - Use 32-bit registers exclusively: Ditch
axand useeax/ecx/edxfor all operations to avoid 16-bit overhead. - Rearrange instructions: Move independent operations (like
inc esi/inc edi) earlier in the loop if possible, so they can run in parallel with memory lookups.
内容的提问来源于stack exchange,提问作者Napoleon

