如何用ARM v7 NEON内联函数获取Q寄存器(int64x2_t)的绝对值?
How to Implement
int64x2_t Absolute Value for ARMv7 NEON Great question! I ran into this exact issue when porting AArch64 NEON code to ARMv7 a while back. Since ARMv7 NEON doesn't have a dedicated vabsq_s64 intrinsic (that's exclusive to AArch64), we can build the 64-bit signed absolute value operation using basic, ARMv7-compatible NEON intrinsics instead.
The Approach
For a 64-bit signed integer, the absolute value can be calculated using bitwise operations:
- For positive numbers: keep the value as-is
- For negative numbers: invert all bits and add 1 (two's complement negation)
We can vectorize this logic using ARMv7 NEON intrinsics:
- Generate a sign mask: Use arithmetic right shift by 63 bits to get a mask where each 64-bit element is
-1(all 1s) if the original value was negative, or0if positive. - Flip bits with XOR: XOR the original vector with the sign mask—this inverts all bits of negative values, leaving positive values unchanged.
- Adjust with subtraction: Subtract the sign mask from the XOR result. For negative values, this is equivalent to adding 1 (since subtracting
-1is the same as adding 1), completing the two's complement negation. For positive values, subtracting0does nothing.
Code Implementation
#include <arm_neon.h> // ARMv7-compatible implementation of int64x2_t absolute value int64x2_t vabsq_s64_v7(int64x2_t a) { // Create sign mask: -1 for negative elements, 0 for positive int64x2_t sign_mask = vshrq_n_s64(a, 63); // Flip bits of negative elements int64x2_t flipped = veorq_s64(a, sign_mask); // Adjust to get absolute value return vsubq_s64(flipped, sign_mask); }
Test Example
You can verify this works with a quick test:
int main() { // Initialize vector with [1, -1] int64x2_t input = vcombine_s64(vcreate_s64(1LL), vcreate_s64(-1LL)); // Compute absolute value int64x2_t abs_result = vabsq_s64_v7(input); // Extract elements to check (optional) int64_t elem0 = vgetq_lane_s64(abs_result, 0); // Should be 1 int64_t elem1 = vgetq_lane_s64(abs_result, 1); // Should be 1 return 0; }
Why This Works
All intrinsics used here are fully supported on ARMv7 NEON:
vshrq_n_s64: Arithmetic right shift of a 64-bit vector by a fixed number of bitsveorq_s64: Vector XOR operationvsubq_s64: Vector subtraction operation
This implementation is efficient, as it only uses three NEON instructions—no branches or loops needed.
内容的提问来源于stack exchange,提问作者Bob liao
相关产品推荐
相关产品推荐

