求_mm256_cvtepi16_epi32的逆操作:AVX2及更早指令集实现方案
Great question! The inverse operation of _mm256_cvtepi16_epi32 (which zero-extends 8x16-bit integers into an 8x32-bit 256-bit vector) doesn't have a single direct AVX2 intrinsic. But we can implement it efficiently with a small sequence of existing instructions, tailored to whether you need saturation or raw truncation:
Option 1: Unsigned Truncation (with Saturation)
If your 32-bit values are unsigned, or you want values exceeding the 16-bit unsigned range (0-65535) to cap at 65535, use this:
#include <immintrin.h> __m128i convert_256i32_to_128i16_unsigned(__m256i vec_32) { // Split the 256-bit vector into two 128-bit chunks (each holding 4x32-bit ints) __m128i low_chunk = _mm256_extracti128_si256(vec_32, 0); __m128i high_chunk = _mm256_extracti128_si256(vec_32, 1); // Pack both chunks into an 8x16-bit 128-bit vector (unsigned saturation) return _mm_packus_epi32(low_chunk, high_chunk); }
Option 2: Signed Truncation (with Saturation)
For signed 32-bit integers, where you want values outside the 16-bit signed range (-32768 to 32767) to saturate to those bounds, use the signed variant:
#include <immintrin.h> __m128i convert_256i32_to_128i16_signed(__m256i vec_32) { __m128i low_chunk = _mm256_extracti128_si256(vec_32, 0); __m128i high_chunk = _mm256_extracti128_si256(vec_32, 1); // Signed saturation pack return _mm_packss_epi32(low_chunk, high_chunk); }
Option 3: Raw Truncation (No Saturation)
If you just want to discard the high 16 bits entirely (ignoring saturation, even if it creates "invalid" 16-bit values for out-of-range inputs), mask first to ensure the pack operation acts as a pure truncation:
#include <immintrin.h> __m128i convert_256i32_to_128i16_truncate(__m256i vec_32) { // Mask off the high 16 bits of each 32-bit element const __m256i mask = _mm256_set1_epi32(0xFFFF); __m256i masked_vec = _mm256_and_si256(vec_32, mask); __m128i low_chunk = _mm256_extracti128_si256(masked_vec, 0); __m128i high_chunk = _mm256_extracti128_si256(masked_vec, 1); // Pack will now just truncate since all values fit in 16 bits return _mm_packus_epi32(low_chunk, high_chunk); }
Performance Notes
This is about as efficient as it gets on AVX2:
_mm256_extracti128_si256is essentially free (it's just a register rename, no execution unit needed)- The
packinstructions are single-cycle, so the entire sequence takes just 2-3 cycles total.
If you had AVX-512, there's a direct _mm256_cvtepi32_epi16 intrinsic, but since you're limited to AVX2 or earlier, the above methods are the optimal approach.
内容的提问来源于stack exchange,提问作者Steve Burns

