如何高效在AVX2通道间重排交错8位值(灰度转BGRA场景)
灰度图转BGRA的AVX2实现:解决跨通道重排问题
我正在用AVX2实现16位窄带灰度图转BGRA格式的转换。这个灰度格式比较特殊:每个像素用16位存储,但实际有效灰度值只在低8位,高8位无意义,本质是8位数据放在16位空间里。
常规C++实现
auto * s = reinterpret_cast<uint16_t*>(input_data); auto * d = reinterpret_cast<uint8_t *>(output_data); for (auto y = 0; y < Height; y++, s += input_pitch, d += output_pitch) { for (auto x = 0; x < Width; x++) { auto v = static_cast<uint8_t>(s[x]); d[x * 4 + 0] = v; // B通道 d[x * 4 + 1] = v; // G通道 d[x * 4 + 2] = v; // R通道 d[x * 4 + 3] = 255;// Alpha通道 } }
AVX2实现思路与问题
我计划用两个__m256i寄存器加载连续32个像素(每个寄存器存16个16位字),然后通过窄化操作丢弃高8位,代码如下:
// 从源加载32个灰度像素(每个像素2字节) __m256i a = _mm256_load_si256(reinterpret_cast<const __m256i*>(s + x)); __m256i b = _mm256_load_si256(reinterpret_cast<const __m256i*>(s + x + 16)); // 将16位灰度值打包为8位 __m256i packed = _mm256_packus_epi16(a, b);
但_mm256_packus_epi16的输出顺序不符合需求:它会把结果排列成a0~a7, b0~b7, a8~a15, b8~b15,而我需要的是a0~a15, b0~b15的顺序。
AVX2没有直接的跨通道置换指令,需要组合多条指令来实现重排。我写了一个测试程序调试,但重排步骤一直失败,希望得到帮助:
测试程序
#include <immintrin.h> #include <iostream> #include <vector> #include <iomanip> // 辅助函数:以字节形式打印256位寄存器 void print_m256i(char const * label, __m256i reg) { uint8_t vals[32]; _mm256_storeu_si256(reinterpret_cast<__m256i*>(vals), reg); std::cout << label << ": "; for (int i = 0; i < 32; ++i) { std::cout << std::setw(3) << static_cast<int>(vals[i]) << " "; } std::cout << "\n"; } int main(void) { // 模拟16位灰度值 alignas(32) uint16_t s1_vals[16] = { 21, 19, 24, 27, 27, 30, 41, 38, 32, 41, 35, 38, 41, 38, 38, 30 }; alignas(32) uint16_t s2_vals[16] = { 30, 30, 30, 30, 24, 24, 24, 24, 24, 32, 32, 32, 35, 35, 32, 32 }; // 加载源数据到AVX2寄存器 __m256i s1 = _mm256_load_si256(reinterpret_cast<const __m256i*>(s1_vals)); __m256i s2 = _mm256_load_si256(reinterpret_cast<const __m256i*>(s2_vals)); // 步骤1:将16位值打包为8位 __m256i packed = _mm256_packus_epi16(s1, s2); print_m256i("Packed", packed); // 步骤2:拆分为两个128位通道以便重排 __m128i lower_lane = _mm256_castsi256_si128(packed); __m128i upper_lane = _mm256_extracti128_si256(packed, 1); print_m256i("Lower Lane", _mm256_castsi128_si256(lower_lane)); print_m256i("Upper Lane", _mm256_castsi128_si256(upper_lane)); // 步骤3:通道内洗牌(调试用,当前逻辑不正确) __m128i interleaved_s1 = _mm_unpacklo_epi8(lower_lane, upper_lane); // 预期s1[0..7]和s1[8..15] __m128i interleaved_s2 = _mm_unpackhi_epi8(lower_lane, upper_lane); // 预期s2[0..7]和s2[8..15] print_m256i("Interleaved S1", _mm256_castsi128_si256(interleaved_s1)); print_m256i("Interleaved S2", _mm256_castsi128_si256(interleaved_s2)); // 步骤4:合并回256位寄存器 __m256i reordered = _mm256_set_m128i(interleaved_s2, interleaved_s1); print_m256i("Reordered", reordered); return 0; }
内容的提问来源于stack exchange,提问作者Robinson
相关产品推荐
相关产品推荐

