x86-64汇编为何采用寄存器而非栈传递函数参数?
Hey there! Great question—this is a super common confusion when switching from 32-bit x86 to x86-64 assembly, since the two use totally different calling conventions. Let's break this down using your code as an example.
First, the key context: x86-64 Calling Conventions
Unlike 32-bit x86 (where nearly all function arguments are passed on the stack), x86-64 follows the System V AMD64 ABI (the standard used on Linux, macOS, and most Unix-like systems). This convention prioritizes passing the first 6 integer/pointer arguments in registers to boost performance:
- 1st argument:
rdi - 2nd argument:
rsi - 3rd argument:
rdx - 4th argument:
rcx - 5th argument:
r8 - 6th argument:
r9
Only arguments beyond the 6th get pushed onto the stack.
Why registers instead of the stack?
The main reason is speed:
- Registers live directly on the CPU core, so accessing them has far lower latency than reading/writing to the stack (which is in system memory). Cutting down on memory operations makes your code run faster.
- It eliminates the overhead of pushing arguments onto the stack before a function call, and popping them off afterward—saving valuable CPU cycles.
What's with the stack storage in your add function?
Looking at your add assembly, you'll see the function copies edi, esi, and edx into the stack frame ([rbp-20], [rbp-24], [rbp-28]). This ties into your note about edx being a volatile register:
- Volatile (caller-saved) registers like
rdi,rsi,rdxcan be modified by the called function without any obligation to restore their original values. - Since your
addfunction needs to reuse the argument values for addition, the compiler (running with-O0, no optimizations) saves them to the stack first. This ensures the values don't get overwritten if other instructions use those registers later in the function. - If you enable optimizations (e.g.,
-O1), the compiler will skip this stack storage and work directly with the registers—your assembly will become much shorter!
When does the stack get used for arguments?
If your function had 7 or more integer arguments, the 7th and beyond would be pushed onto the stack before calling the function. The register-only approach is just an optimization for the common case of small argument lists.
To recap: x86-64 doesn't abandon stack-based argument passing entirely—it just uses registers first because they're faster. The choice depends on the calling convention, which evolved to take advantage of x86-64's larger set of general-purpose registers compared to 32-bit x86.
内容的提问来源于stack exchange,提问作者killerprince182

