如何实现3x3复合卷积核与像素数组图像的卷积运算(SIMD优化)
基于SIMD的复合卷积实现补充
首先,我们可以先简化复合卷积的计算逻辑,这能大幅降低SIMD实现的复杂度。先把卷积核H和V的结果相加:
H卷积结果(每个输出像素):
src[y-1][x-1] + src[y-1][x+1] - src[y+1][x-1] - src[y+1][x+1]
V卷积结果(每个输出像素):
src[y-1][x-1] - src[y-1][x+1] + src[y+1][x-1] - src[y+1][x+1]
将两者相加后,中间项会抵消,最终得到简化后的复合卷积公式:
output[x][y] = 2 * src[y-1][x-1] - 2 * src[y+1][x+1]
这个简化是关键——我们只需要用到上一行的当前列像素,以及下一行的当前列+2的像素,完全不需要中间行的任何数据。
接下来,我们基于这个公式补充SIMD实现代码:
补充前8个像素的卷积逻辑
在你标记的//there is we make the first convolution for 8px's位置,插入以下代码:
// 准备零向量用于后续打包操作 __m128i vec_zero = _mm_setzero_si128(); // 提取下一行中对应输出像素的x+1位置的像素:src[y+1][2] ~ src[y+1][9] // 左移2字节后,低8字节正好是原str3_16pxs的第2到第9个像素 __m128i str3_bottom_right_1st8 = _mm_slli_si128(str3_16pxs, 2); // 扩展为16位整数,避免计算溢出 str3_bottom_right_1st8 = _mm_cvtepu8_epi16(str3_bottom_right_1st8); // 计算2*上一行像素 - 2*下一行右移2位的像素 // 左移1位等价于乘2,用sub实现减法 __m128i conv_result_1st8 = _mm_sub_epi16( _mm_slli_epi16(str1_16pxs_pack1st_8to16, 1), _mm_slli_epi16(str3_bottom_right_1st8, 1) ); // 将16位结果饱和转换为8位无符号像素(避免溢出到0-255范围外) __m128i conv_result_1st8_packed = _mm_packus_epi16(conv_result_1st8, vec_zero); // 存储前8个输出像素(低8字节) _mm_storel_epi64((__m128i*)dst_all_str, conv_result_1st8_packed); dst_all_str += 8;
补充剩余8个像素的卷积逻辑
在你标记的//do convolution for the remaining 8px's and so on until the end of the read line位置,插入以下代码:
// 提取下一行中对应剩余8个输出像素的x+1位置:src[y+1][10] ~ src[y+1][17] // 先右移8字节取后8个像素,再左移2字节得到目标位置 __m128i str3_bottom_right_2nd8 = _mm_slli_si128(str3_16pxs_plus_8pxs, 2); str3_bottom_right_2nd8 = _mm_cvtepu8_epi16(str3_bottom_right_2nd8); // 用同样的简化公式计算卷积 __m128i conv_result_2nd8 = _mm_sub_epi16( _mm_slli_epi16(str1_16pxs_pack2nd_8to16, 1), _mm_slli_epi16(str3_bottom_right_2nd8, 1) ); // 饱和转换为8位并存储 __m128i conv_result_2nd8_packed = _mm_packus_epi16(conv_result_2nd8, vec_zero); _mm_storel_epi64((__m128i*)dst_all_str, conv_result_2nd8_packed); dst_all_str += 8;
关键细节说明
- 饱和转换:使用
_mm_packus_epi16确保卷积结果超出0-255范围时,自动截断到合法的像素值范围,避免出现异常颜色。 - 移位代替乘法:用
_mm_slli_epi16(..., 1)代替乘法指令,因为乘2等价于左移1位,性能更高。 - 向量提取:通过
_mm_slli_si128实现字节级的移位,快速定位到下一行中需要的x+1位置的像素,避免额外的内存加载。
内容的提问来源于stack exchange,提问作者Georgy
相关产品推荐
相关产品推荐

