You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求_mm256_cvtepi16_epi32的逆操作:AVX2及更早指令集实现方案

Great question! The inverse operation of _mm256_cvtepi16_epi32 (which zero-extends 8x16-bit integers into an 8x32-bit 256-bit vector) doesn't have a single direct AVX2 intrinsic. But we can implement it efficiently with a small sequence of existing instructions, tailored to whether you need saturation or raw truncation:

Option 1: Unsigned Truncation (with Saturation)

If your 32-bit values are unsigned, or you want values exceeding the 16-bit unsigned range (0-65535) to cap at 65535, use this:

#include <immintrin.h>

__m128i convert_256i32_to_128i16_unsigned(__m256i vec_32) {
    // Split the 256-bit vector into two 128-bit chunks (each holding 4x32-bit ints)
    __m128i low_chunk = _mm256_extracti128_si256(vec_32, 0);
    __m128i high_chunk = _mm256_extracti128_si256(vec_32, 1);
    // Pack both chunks into an 8x16-bit 128-bit vector (unsigned saturation)
    return _mm_packus_epi32(low_chunk, high_chunk);
}

Option 2: Signed Truncation (with Saturation)

For signed 32-bit integers, where you want values outside the 16-bit signed range (-32768 to 32767) to saturate to those bounds, use the signed variant:

#include <immintrin.h>

__m128i convert_256i32_to_128i16_signed(__m256i vec_32) {
    __m128i low_chunk = _mm256_extracti128_si256(vec_32, 0);
    __m128i high_chunk = _mm256_extracti128_si256(vec_32, 1);
    // Signed saturation pack
    return _mm_packss_epi32(low_chunk, high_chunk);
}

Option 3: Raw Truncation (No Saturation)

If you just want to discard the high 16 bits entirely (ignoring saturation, even if it creates "invalid" 16-bit values for out-of-range inputs), mask first to ensure the pack operation acts as a pure truncation:

#include <immintrin.h>

__m128i convert_256i32_to_128i16_truncate(__m256i vec_32) {
    // Mask off the high 16 bits of each 32-bit element
    const __m256i mask = _mm256_set1_epi32(0xFFFF);
    __m256i masked_vec = _mm256_and_si256(vec_32, mask);
    
    __m128i low_chunk = _mm256_extracti128_si256(masked_vec, 0);
    __m128i high_chunk = _mm256_extracti128_si256(masked_vec, 1);
    // Pack will now just truncate since all values fit in 16 bits
    return _mm_packus_epi32(low_chunk, high_chunk);
}

Performance Notes

This is about as efficient as it gets on AVX2:

  • _mm256_extracti128_si256 is essentially free (it's just a register rename, no execution unit needed)
  • The pack instructions are single-cycle, so the entire sequence takes just 2-3 cycles total.

If you had AVX-512, there's a direct _mm256_cvtepi32_epi16 intrinsic, but since you're limited to AVX2 or earlier, the above methods are the optimal approach.

内容的提问来源于stack exchange,提问作者Steve Burns

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:22:44