ARM子程序调用与链接寄存器(LR)的必要性探究:栈帧可存储返回地址场景下的疑问
Great question! Let’s break down why LR is a critical part of ARM’s design, even though stack frames already handle return address storage for stack backtracking. Here are the key reasons:
Cut down on stack overhead for function calls
When you use theBL(Branch with Link) instruction to call a function, ARM automatically saves the return address (the next instruction that should run after the function returns) into LR. Without LR, every single function call would require an extra explicit instruction to push the PC onto the stack—likeSTR PC, [SP, #-4]!. For nested calls (think A calling B, which calls C), LR lets each function hold its own return address temporarily, avoiding repeated stack pushes and pops that would slow things down and use more memory bandwidth.Streamline interrupt and exception handling
ARM’s exception modes (IRQ, FIQ, etc.) have their own dedicated LR registers. When an interrupt fires, the hardware automatically loads the correct return address into the mode-specific LR. This removes the need for manual stack operations to save the return address when entering an interrupt, making exception handlers faster and less error-prone. When exiting the interrupt, a simple instruction likeSUBS PC, LR, #4(adjusting for the pipeline delay) jumps you straight back to the interrupted code.Enable leaf function optimizations
Leaf functions—those that don’t call any other functions—can skip pushing LR onto the stack entirely. Since they never overwrite LR, they can return directly usingBX LR(orMOV PC, LRfor older ARM cores) without touching the stack at all. This saves stack space and the cycle cost of push/pop operations, which is a big win for small, frequently used utility functions.Simplify return sequences and reduce errors
Even for non-leaf functions, LR makes returns cleaner. Instead of having to load the return address directly from the stack, you can pop LR first then useBX LR. More importantly, theBLinstruction handles calculating the correct return address (which is PC+4, thanks to ARM’s pipeline) automatically and stores it in LR—so you don’t have to manually compute it in assembly, reducing the chance of bugs.
内容的提问来源于stack exchange,提问作者Srinivas

