如何用AARCH64汇编实现符合特定约束的memcpy函数?
Got it, let's build this AARCH64 memcpy variant that fits all your constraints. Since we're guaranteed aligned addresses, no overlapping memory, and a length that's a multiple of 16, we can optimize for speed while keeping the code tight and under that 80-byte limit.
First, here's an optimized implementation that handles 32 bytes per loop iteration (double the 16-byte block size) to minimize loop overhead, while adhering to all rules:
.global memcpy_aligned memcpy_aligned: // AARCH64 calling convention: // x0 = destination address (16-byte aligned) // x1 = source address (16-byte aligned) // x2 = number of bytes to copy (multiple of 16) // Returns: initial destination address (matches standard memcpy) mov x3, x0 // Save the original dest address for return cbz x2, .Lexit // Early exit if there's nothing to copy lsr x2, x2, #4 // Convert byte count to 16-byte block count // (e.g., 64 bytes → 4 blocks) .Lloop: ldp q0, q1, [x1], #32 // Load two 16-byte NEON blocks (32 bytes total) // from src, then increment src by 32 stp q0, q1, [x0], #32 // Store the two blocks to dest, then increment dest by 32 subs x2, x2, #2 // Decrement block count by 2 (we handled two blocks) bne .Lloop // Loop until all blocks are processed .Lexit: mov x0, x3 // Restore original dest address as return value ret // Return to caller
Breakdown of the Code
Let's walk through each line to clarify what's happening:
.global memcpy_aligned: Makes the function visible to other files (like C code that calls it).mov x3, x0: Saves the initial destination address because we'll incrementx0during copying, and standard memcpy returns the original pointer.cbz x2, .Lexit: Checks if the byte count is zero—if so, skip the loop entirely for efficiency.lsr x2, x2, #4: Right-shifts the byte count by 4 bits (equivalent to dividing by 16) to get the number of 16-byte blocks we need to copy..Lloop: The start of our copy loop.ldp q0, q1, [x1], #32: Uses the Load Pair instruction to grab two 128-bit (16-byte) NEON registers (q0,q1) from the source address. The#32post-incrementsx1by 32 bytes after the load.stp q0, q1, [x0], #32: Uses Store Pair to write the two NEON registers to the destination address, then incrementsx0by 32 bytes.subs x2, x2, #2: Subtracts 2 from the block count (since we handled two 16-byte blocks) and sets the CPU's status flags.bne .Lloop: If the block count isn't zero (flags indicate "not equal"), jump back to the start of the loop.mov x0, x3: Restores the original destination address tox0(the register used for return values in AARCH64).ret: Returns control to the caller.
Why This Fits Your Constraints
- 16-byte alignment:
ldp/stpwork optimally (and safely) with aligned addresses, which your inputs guarantee. Since we're moving 32 bytes at a time (a multiple of 16), address alignment is maintained throughout the loop. - No overlapping memory: We don't need to handle reverse copying (required for overlapping regions) because your assumption rules this out—so incrementing addresses forward is safe.
- Length is 16x: The
lsr x2, x2, #4converts the byte count to an integer number of 16-byte blocks, and we decrement by 2 each loop (handling two blocks) until we hit zero. No leftover bytes to worry about. - Instruction length: Each AARCH64 instruction is 4 bytes. This code has 11 instructions, totaling 44 bytes—way under the 80-byte limit.
Simpler (Single Block) Alternative
If you prefer a more straightforward implementation that handles one 16-byte block per iteration, here's a version that's even shorter:
.global memcpy_aligned_simple memcpy_aligned_simple: mov x3, x0 cbz x2, .Lexit_simple lsr x2, x2, #4 .Lloop_simple: ld1 {v0.16b}, [x1], #16 // Load one 16-byte block into NEON register v0 st1 {v0.16b}, [x0], #16 // Store the block to dest subs x2, x2, #1 bne .Lloop_simple .Lexit_simple: mov x0, x3 ret
This version uses ld1/st1 (single-element load/store for NEON) and processes one 16-byte block per loop. It's 10 instructions (40 bytes) and still meets all your requirements.
内容的提问来源于stack exchange,提问作者Quadratic Ocelot

