You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用ARM v7 NEON内联函数获取Q寄存器(int64x2_t)的绝对值?

How to Implement int64x2_t Absolute Value for ARMv7 NEON

Great question! I ran into this exact issue when porting AArch64 NEON code to ARMv7 a while back. Since ARMv7 NEON doesn't have a dedicated vabsq_s64 intrinsic (that's exclusive to AArch64), we can build the 64-bit signed absolute value operation using basic, ARMv7-compatible NEON intrinsics instead.

The Approach

For a 64-bit signed integer, the absolute value can be calculated using bitwise operations:

  • For positive numbers: keep the value as-is
  • For negative numbers: invert all bits and add 1 (two's complement negation)

We can vectorize this logic using ARMv7 NEON intrinsics:

  1. Generate a sign mask: Use arithmetic right shift by 63 bits to get a mask where each 64-bit element is -1 (all 1s) if the original value was negative, or 0 if positive.
  2. Flip bits with XOR: XOR the original vector with the sign mask—this inverts all bits of negative values, leaving positive values unchanged.
  3. Adjust with subtraction: Subtract the sign mask from the XOR result. For negative values, this is equivalent to adding 1 (since subtracting -1 is the same as adding 1), completing the two's complement negation. For positive values, subtracting 0 does nothing.

Code Implementation

#include <arm_neon.h>

// ARMv7-compatible implementation of int64x2_t absolute value
int64x2_t vabsq_s64_v7(int64x2_t a) {
    // Create sign mask: -1 for negative elements, 0 for positive
    int64x2_t sign_mask = vshrq_n_s64(a, 63);
    // Flip bits of negative elements
    int64x2_t flipped = veorq_s64(a, sign_mask);
    // Adjust to get absolute value
    return vsubq_s64(flipped, sign_mask);
}

Test Example

You can verify this works with a quick test:

int main() {
    // Initialize vector with [1, -1]
    int64x2_t input = vcombine_s64(vcreate_s64(1LL), vcreate_s64(-1LL));
    // Compute absolute value
    int64x2_t abs_result = vabsq_s64_v7(input);
    
    // Extract elements to check (optional)
    int64_t elem0 = vgetq_lane_s64(abs_result, 0); // Should be 1
    int64_t elem1 = vgetq_lane_s64(abs_result, 1); // Should be 1
    
    return 0;
}

Why This Works

All intrinsics used here are fully supported on ARMv7 NEON:

  • vshrq_n_s64: Arithmetic right shift of a 64-bit vector by a fixed number of bits
  • veorq_s64: Vector XOR operation
  • vsubq_s64: Vector subtraction operation

This implementation is efficient, as it only uses three NEON instructions—no branches or loops needed.

内容的提问来源于stack exchange,提问作者Bob liao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 10:52:37