You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现3x3复合卷积核与像素数组图像的卷积运算(SIMD优化)

基于SIMD的复合卷积实现补充

首先,我们可以先简化复合卷积的计算逻辑,这能大幅降低SIMD实现的复杂度。先把卷积核H和V的结果相加:

H卷积结果(每个输出像素):

src[y-1][x-1] + src[y-1][x+1] - src[y+1][x-1] - src[y+1][x+1]

V卷积结果(每个输出像素):

src[y-1][x-1] - src[y-1][x+1] + src[y+1][x-1] - src[y+1][x+1]

将两者相加后,中间项会抵消,最终得到简化后的复合卷积公式:

output[x][y] = 2 * src[y-1][x-1] - 2 * src[y+1][x+1]

这个简化是关键——我们只需要用到上一行的当前列像素,以及下一行的当前列+2的像素,完全不需要中间行的任何数据。

接下来,我们基于这个公式补充SIMD实现代码:

补充前8个像素的卷积逻辑

在你标记的//there is we make the first convolution for 8px's位置,插入以下代码:

// 准备零向量用于后续打包操作
__m128i vec_zero = _mm_setzero_si128();

// 提取下一行中对应输出像素的x+1位置的像素:src[y+1][2] ~ src[y+1][9]
// 左移2字节后,低8字节正好是原str3_16pxs的第2到第9个像素
__m128i str3_bottom_right_1st8 = _mm_slli_si128(str3_16pxs, 2);
// 扩展为16位整数,避免计算溢出
str3_bottom_right_1st8 = _mm_cvtepu8_epi16(str3_bottom_right_1st8);

// 计算2*上一行像素 - 2*下一行右移2位的像素
// 左移1位等价于乘2,用sub实现减法
__m128i conv_result_1st8 = _mm_sub_epi16(
    _mm_slli_epi16(str1_16pxs_pack1st_8to16, 1),
    _mm_slli_epi16(str3_bottom_right_1st8, 1)
);

// 将16位结果饱和转换为8位无符号像素(避免溢出到0-255范围外)
__m128i conv_result_1st8_packed = _mm_packus_epi16(conv_result_1st8, vec_zero);
// 存储前8个输出像素(低8字节)
_mm_storel_epi64((__m128i*)dst_all_str, conv_result_1st8_packed);
dst_all_str += 8;

补充剩余8个像素的卷积逻辑

在你标记的//do convolution for the remaining 8px's and so on until the end of the read line位置,插入以下代码:

// 提取下一行中对应剩余8个输出像素的x+1位置:src[y+1][10] ~ src[y+1][17]
// 先右移8字节取后8个像素,再左移2字节得到目标位置
__m128i str3_bottom_right_2nd8 = _mm_slli_si128(str3_16pxs_plus_8pxs, 2);
str3_bottom_right_2nd8 = _mm_cvtepu8_epi16(str3_bottom_right_2nd8);

// 用同样的简化公式计算卷积
__m128i conv_result_2nd8 = _mm_sub_epi16(
    _mm_slli_epi16(str1_16pxs_pack2nd_8to16, 1),
    _mm_slli_epi16(str3_bottom_right_2nd8, 1)
);

// 饱和转换为8位并存储
__m128i conv_result_2nd8_packed = _mm_packus_epi16(conv_result_2nd8, vec_zero);
_mm_storel_epi64((__m128i*)dst_all_str, conv_result_2nd8_packed);
dst_all_str += 8;

关键细节说明

  • 饱和转换:使用_mm_packus_epi16确保卷积结果超出0-255范围时,自动截断到合法的像素值范围,避免出现异常颜色。
  • 移位代替乘法:用_mm_slli_epi16(..., 1)代替乘法指令,因为乘2等价于左移1位,性能更高。
  • 向量提取:通过_mm_slli_si128实现字节级的移位,快速定位到下一行中需要的x+1位置的像素,避免额外的内存加载。

内容的提问来源于stack exchange,提问作者Georgy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:41:21