为何8位MUL指令结果存入AX寄存器,而16位、32位MUL指令结果拆分存入[E]DX:[E]AX?
Great question—this dives into the historical design roots of x86 and how register conventions evolved to maintain consistency across instruction widths. Let’s break this down step by step:
1. 8086’s Accumulator-Centric Architecture
The original 8086 was built around a compact register set where AX (and its 8-bit sub-registers AL/AH) served as the primary accumulator for arithmetic operations. For 8-bit MUL CL, the implicit source operand is always AL—this choice kept instruction encoding small (no need to encode both source registers). Since multiplying two 8-bit values produces a 16-bit result, storing it in AX made perfect sense:
- It aligned with the register hierarchy (AL is the low byte of AX, AH the high byte). Using DL:AL would have crossed disjoint register pairs (DL belongs to DX, not AX), breaking the intuitive grouping of 8-bit sub-registers under their 16-bit parents.
- AX was already the go-to register for holding arithmetic results, so programmers could immediately use the 16-bit result in subsequent operations without combining two separate 8-bit registers.
2. Consistency Across MUL Widths
The pattern scales logically for wider MUL instructions, using register pairs only when the result exceeds the width of a single register:
- 16-bit
MUL BX: Two 16-bit values multiply to a 32-bit result. The 8086 had no single 32-bit register, so the result splits across DX:AX (DX holds the high 16 bits, AX the low). - 32-bit
MUL EBX: Two 32-bit values produce a 64-bit result, which fits in the EDX:EAX pair (standard for 32-bit x86). - 64-bit
MUL RCX: Two 64-bit values multiply to a 128-bit result. x86-64 has no single 128-bit general-purpose register (vector registers like XMM are separate), so we use RDX:RAX to hold the full result.
A quick correction to your note: 16-bit MUL’s 32-bit result can’t fit in EAX alone (even in 32-bit mode, it uses DX:AX, zero-extended to EDX:EAX). Similarly, 32-bit MUL’s 64-bit result needs EDX:EAX, not just RAX. The core rule is: n-bit MUL produces a 2n-bit result, stored in the widest available register(s) that fit—single register if 2n matches an existing register width, pair if not.
For 8-bit MUL, 2*8=16 bits exactly matches AX’s width—so there’s no need to split across two 8-bit registers like DL:AL. This kept the instruction set consistent and intuitive for programmers already familiar with AX as the main accumulator.
内容的提问来源于stack exchange,提问作者user15177006

