ARM64架构ISR中高效压入全部寄存器至栈的优化问询
Great question—since your interrupt fires extremely often, every cycle saved adds up quickly. Your current code works, but we can optimize it by reducing redundant stack pointer operations and improving memory access patterns. Let's break down the key improvements:
1. Minimize Stack Pointer (SP) Modifications
Your current code uses stp xn, xm, [sp, #-16]! and str xn, [sp, #-16]! repeatedly, which modifies the stack pointer every single time. Each SP write has a small overhead, and with 34 total SP updates in your save sequence, this adds up fast for frequent interrupts.
Instead, calculate the total stack space needed upfront, adjust SP once, then use fixed-offset stores/loads. This cuts SP modifications from 34 to just 2 (one for saving, one for restoring).
2. Optimized Save/Restore Sequence
Here's a revised implementation that applies this approach. We also ensure proper 16-byte stack alignment required by ARM64:
Save Context
; Calculate total stack size: ; 31 GP registers (x0-x30) = 31*8 = 248 bytes ; 32 SIMD registers (q0-q31) = 32*16 = 512 bytes ; FPCR + FPSR = 2*8 = 16 bytes ; Total: 248 + 512 +16 = 776. Add 8 bytes to align to 16-byte boundary → 784 sub sp, sp, #784 ; Save general-purpose registers (contiguous block) stp x0, x1, [sp, #0] stp x2, x3, [sp, #16] stp x4, x5, [sp, #32] stp x6, x7, [sp, #48] stp x8, x9, [sp, #64] stp x10, x11, [sp, #80] stp x12, x13, [sp, #96] stp x14, x15, [sp, #112] stp x16, x17, [sp, #128] stp x18, x19, [sp, #144] stp x20, x21, [sp, #160] stp x22, x23, [sp, #176] stp x24, x25, [sp, #192] stp x26, x27, [sp, #208] stp x28, x29, [sp, #224] str x30, [sp, #240] ; Save SIMD/FPU registers (contiguous block) stp q0, q1, [sp, #248] stp q2, q3, [sp, #248+32] stp q4, q5, [sp, #248+64] stp q6, q7, [sp, #248+96] stp q8, q9, [sp, #248+128] stp q10, q11, [sp, #248+160] stp q12, q13, [sp, #248+192] stp q14, q15, [sp, #248+224] stp q16, q17, [sp, #248+256] stp q18, q19, [sp, #248+288] stp q20, q21, [sp, #248+320] stp q22, q23, [sp, #248+352] stp q24, q25, [sp, #248+384] stp q26, q27, [sp, #248+416] stp q28, q29, [sp, #248+448] stp q30, q31, [sp, #248+480] ; Save FPCR and FPSR mrs x0, FPCR mrs x1, FPSR str x0, [sp, #248+512] str x1, [sp, #248+512+8] bl vector_irq
Restore Context
; Restore FPCR and FPSR ldr x0, [sp, #248+512] ldr x1, [sp, #248+512+8] msr FPCR, x0 msr FPSR, x1 ; Restore SIMD/FPU registers (reverse order to match save) ldp q30, q31, [sp, #248+480] ldp q28, q29, [sp, #248+448] ldp q26, q27, [sp, #248+416] ldp q24, q25, [sp, #248+384] ldp q22, q23, [sp, #248+352] ldp q20, q21, [sp, #248+320] ldp q18, q19, [sp, #248+288] ldp q16, q17, [sp, #248+256] ldp q14, q15, [sp, #248+224] ldp q12, q13, [sp, #248+192] ldp q10, q11, [sp, #248+160] ldp q8, q9, [sp, #248+128] ldp q6, q7, [sp, #248+96] ldp q4, q5, [sp, #248+64] ldp q2, q3, [sp, #248+32] ldp q0, q1, [sp, #248] ; Restore general-purpose registers (reverse order to match save) ldr x30, [sp, #240] ldp x28, x29, [sp, #224] ldp x26, x27, [sp, #208] ldp x24, x25, [sp, #192] ldp x22, x23, [sp, #176] ldp x20, x21, [sp, #160] ldp x18, x19, [sp, #144] ldp x16, x17, [sp, #128] ldp x14, x15, [sp, #112] ldp x12, x13, [sp, #96] ldp x10, x11, [sp, #80] ldp x8, x9, [sp, #64] ldp x6, x7, [sp, #48] ldp x4, x5, [sp, #32] ldp x2, x3, [sp, #16] ldp x0, x1, [sp, #0] ; Restore stack pointer add sp, sp, #784 eret
3. Additional Optimizations to Consider
- Skip SIMD Registers If Unused: If your system never uses SIMD/FPU instructions (e.g., bare-metal with no floating-point code), you can completely omit saving q0-q31, FPCR, and FPSR. This cuts the save/restore sequence in half—huge savings for frequent interrupts. Just note: if interrupts can occur while SIMD code is running, you can't skip this (you'll corrupt the interrupted context).
- Leverage CPU Pipeline Parallelism: ARM64 CPUs execute instructions out of order when possible. Grouping
stp/ldpinstructions that access non-overlapping memory regions helps the CPU optimize throughput, which the contiguous block approach above already supports. - Hardware-Assisted Context Switching: Some specialized ARM64 chips include custom instructions for fast context saving (e.g.,
savectx). Check your platform's documentation—if available, this can be even faster than manual register saves.
Why This Works Better
By reducing SP modifications from 34 to 2, we eliminate redundant write operations to a frequently accessed register. The contiguous memory blocks also improve cache locality, making the save/restore operations more predictable for the CPU's memory pipeline. For interrupts that fire thousands of times per second, these small cycle savings add up to measurable performance gains.
内容的提问来源于stack exchange,提问作者qwerty123443

