You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于ARM Neon高效统计16字节缓冲区不同值数量及相关判断

Absolutely! ARM Neon's SIMD instructions are perfect for optimizing this kind of fixed-size (16-byte) buffer operation—both for counting distinct values and checking common repetition patterns like all elements being identical, exactly two distinct values, or more. Let's break this down step by step.

1. Optimized Distinct Value Count with Neon

The original scalar algorithm uses a 256-byte lookup array to track seen values, which works but misses out on SIMD efficiencies. With Neon, we can use a bitmask approach (since uint8_t values range from 0-255, we map each value to a single bit in a 256-bit mask) to track seen values, then count the set bits for the distinct count. This is far more cache-efficient and leverages Neon's specialized bit-counting instructions.

Here's the implementation using Neon intrinsics:

#include <arm_neon.h>

unsigned neon_getCount(const uint8_t data[16]) {
    // Load the 16-byte buffer into a Neon vector
    uint8x16_t data_vec = vld1q_u8(data);
    
    // Initialize 256-bit mask: low 128 bits for values 0-127, high 128 for 128-255
    uint64x2_t mask_low = vdupq_n_u64(0);
    uint64x2_t mask_high = vdupq_n_u64(0);

    // Mark each value's corresponding bit in the mask
    for (int i = 0; i < 16; ++i) {
        uint8_t val = vgetq_lane_u8(data_vec, i);
        if (val < 128) {
            uint64x2_t bit_mask = vdupq_n_u64(1ULL << val);
            mask_low = vorrq_u64(mask_low, bit_mask);
        } else {
            uint64x2_t bit_mask = vdupq_n_u64(1ULL << (val - 128));
            mask_high = vorrq_u64(mask_high, bit_mask);
        }
    }

    // Count set bits in each mask half and sum the results
    unsigned count_low = vaddvq_u32(vreinterpretq_u32_u64(vcntq_u64(mask_low)));
    unsigned count_high = vaddvq_u32(vreinterpretq_u32_u64(vcntq_u64(mask_high)));

    return count_low + count_high;
}
  • How this works: We map each uint8 value to a unique bit in a 256-bit mask. After processing all elements, the number of set bits directly equals the number of distinct values. Neon's vcntq_u64 counts set bits per 64-bit lane, and vaddvq_u32 sums those counts efficiently in parallel.

2. Efficient Checks for Repetition Patterns

We can skip full distinct counting for these common cases and use Neon to get answers in just a few instructions:

Check if all elements are identical

The fastest method is to compare the minimum and maximum values in the buffer—if they match, every element is the same:

bool neon_all_same(const uint8_t data[16]) {
    uint8x16_t data_vec = vld1q_u8(data);
    return vminvq_u8(data_vec) == vmaxvq_u8(data_vec);
}
  • vminvq_u8 and vmaxvq_u8 are single Neon instructions that compute the min/max of the 16-element vector, making this operation blazingly fast.

Check if exactly two distinct values exist

First confirm the buffer isn't all identical, then verify every element is one of the first two distinct values found:

bool neon_only_two_distinct(const uint8_t data[16]) {
    if (neon_all_same(data)) return false;

    uint8x16_t data_vec = vld1q_u8(data);
    uint8_t first_val = data[0];
    uint8_t second_val = 0;

    // Find the first value that differs from the first element
    for (int i = 1; i < 16; ++i) {
        if (data[i] != first_val) {
            second_val = data[i];
            break;
        }
    }

    // Check if all elements are either first_val or second_val
    uint8x16_t first_vec = vdupq_n_u8(first_val);
    uint8x16_t second_vec = vdupq_n_u8(second_val);
    uint8x16_t eq_first = vceqq_u8(data_vec, first_vec);
    uint8x16_t eq_second = vceqq_u8(data_vec, second_vec);
    uint8x16_t eq_either = vorrq_u8(eq_first, eq_second);

    // Sum the mask: if all bytes are 0xFF (all matches), sum equals 16*255
    return vaddvq_u8(eq_either) == 16 * 0xFF;
}

Alternatively, you could reuse neon_getCount and check if the return value is 2—but this dedicated check is faster since it avoids full bitmask counting.

Check for more than two distinct values

This is simply the inverse of the above two checks:

bool neon_more_than_two_distinct(const uint8_t data[16]) {
    return !neon_all_same(data) && !neon_only_two_distinct(data);
}

Performance Notes

  • For fixed 16-byte buffers, these Neon implementations will outperform the scalar version significantly, especially in tight loops. The bitmask approach avoids cache misses from the 256-byte lookup array in the original code.
  • All functions use standard Neon intrinsics, compatible with most ARM compilers (GCC, Clang, ARM Compiler).

内容的提问来源于stack exchange,提问作者Pavel P

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:22:15