You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用AARCH64汇编实现符合特定约束的memcpy函数?

AARCH64 Aligned Memcpy Implementation (Constrained)

Got it, let's build this AARCH64 memcpy variant that fits all your constraints. Since we're guaranteed aligned addresses, no overlapping memory, and a length that's a multiple of 16, we can optimize for speed while keeping the code tight and under that 80-byte limit.

First, here's an optimized implementation that handles 32 bytes per loop iteration (double the 16-byte block size) to minimize loop overhead, while adhering to all rules:

.global memcpy_aligned
memcpy_aligned:
    // AARCH64 calling convention:
    // x0 = destination address (16-byte aligned)
    // x1 = source address (16-byte aligned)
    // x2 = number of bytes to copy (multiple of 16)
    // Returns: initial destination address (matches standard memcpy)
    
    mov x3, x0              // Save the original dest address for return
    cbz x2, .Lexit          // Early exit if there's nothing to copy
    
    lsr x2, x2, #4          // Convert byte count to 16-byte block count
                            // (e.g., 64 bytes → 4 blocks)
.Lloop:
    ldp q0, q1, [x1], #32   // Load two 16-byte NEON blocks (32 bytes total)
                            // from src, then increment src by 32
    stp q0, q1, [x0], #32   // Store the two blocks to dest, then increment dest by 32
    subs x2, x2, #2         // Decrement block count by 2 (we handled two blocks)
    bne .Lloop              // Loop until all blocks are processed
    
.Lexit:
    mov x0, x3              // Restore original dest address as return value
    ret                     // Return to caller

Breakdown of the Code

Let's walk through each line to clarify what's happening:

  • .global memcpy_aligned: Makes the function visible to other files (like C code that calls it).
  • mov x3, x0: Saves the initial destination address because we'll increment x0 during copying, and standard memcpy returns the original pointer.
  • cbz x2, .Lexit: Checks if the byte count is zero—if so, skip the loop entirely for efficiency.
  • lsr x2, x2, #4: Right-shifts the byte count by 4 bits (equivalent to dividing by 16) to get the number of 16-byte blocks we need to copy.
  • .Lloop: The start of our copy loop.
  • ldp q0, q1, [x1], #32: Uses the Load Pair instruction to grab two 128-bit (16-byte) NEON registers (q0, q1) from the source address. The #32 post-increments x1 by 32 bytes after the load.
  • stp q0, q1, [x0], #32: Uses Store Pair to write the two NEON registers to the destination address, then increments x0 by 32 bytes.
  • subs x2, x2, #2: Subtracts 2 from the block count (since we handled two 16-byte blocks) and sets the CPU's status flags.
  • bne .Lloop: If the block count isn't zero (flags indicate "not equal"), jump back to the start of the loop.
  • mov x0, x3: Restores the original destination address to x0 (the register used for return values in AARCH64).
  • ret: Returns control to the caller.

Why This Fits Your Constraints

  • 16-byte alignment: ldp/stp work optimally (and safely) with aligned addresses, which your inputs guarantee. Since we're moving 32 bytes at a time (a multiple of 16), address alignment is maintained throughout the loop.
  • No overlapping memory: We don't need to handle reverse copying (required for overlapping regions) because your assumption rules this out—so incrementing addresses forward is safe.
  • Length is 16x: The lsr x2, x2, #4 converts the byte count to an integer number of 16-byte blocks, and we decrement by 2 each loop (handling two blocks) until we hit zero. No leftover bytes to worry about.
  • Instruction length: Each AARCH64 instruction is 4 bytes. This code has 11 instructions, totaling 44 bytes—way under the 80-byte limit.

Simpler (Single Block) Alternative

If you prefer a more straightforward implementation that handles one 16-byte block per iteration, here's a version that's even shorter:

.global memcpy_aligned_simple
memcpy_aligned_simple:
    mov x3, x0
    cbz x2, .Lexit_simple
    lsr x2, x2, #4
.Lloop_simple:
    ld1 {v0.16b}, [x1], #16  // Load one 16-byte block into NEON register v0
    st1 {v0.16b}, [x0], #16  // Store the block to dest
    subs x2, x2, #1
    bne .Lloop_simple
.Lexit_simple:
    mov x0, x3
    ret

This version uses ld1/st1 (single-element load/store for NEON) and processes one 16-byte block per loop. It's 10 instructions (40 bytes) and still meets all your requirements.

内容的提问来源于stack exchange,提问作者Quadratic Ocelot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:08:52