You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用SSE2与C++快速为RGB图像添加Alpha通道?

优化SSE2实现RGB转RGBA的性能问题

我正在使用SSE2编写C++版YUV420p转RGBA的颜色转换算法,目前已实现YUV420p转RGB以及RGB转RGBA功能,测试结果如下:

size of image: 1920 x 1200
time of RGBA to YUV conversion: 0.0029011
time of YUV to RGB conversion: 0.0044585
time of RGB to RGBA conversion (approach 1): 0.0064747
time of RGB to RGBA conversion (approach 2): 0.0066194
time of RGB to RGBA conversion (approach 3): 0.0069835

可见RGB转RGBA的耗时比YUV420p转RGB或RGBA转YUV420p更长。由于在YUV420p转RGB的计算中插入Alpha通道存在诸多困难,我尝试采用后处理步骤(RGB转RGBA),以下是目前的三种实现代码:

方法1

void convertRGB24itoRGBA32i ( int width, int height, const unsigned char *RGB, unsigned char *RGBA ) {
    const size_t numPixels = (width - 1) * (height - 1);

    for ( size_t i = 0; i < numPixels; i++ ) 
    {
        __m128i sourcePixel = _mm_loadu_si128 ( (__m128i*)&RGB[i * 3] );
        //__m128i alphaChannel = _mm_setzero_si128 ( ); // Set alpha to 0 (transparent)
        __m128i alphaChannel = _mm_set1_epi32 ( 0xFF000000 );
        __m128i rgb32Pixel = _mm_or_si128 ( alphaChannel, sourcePixel );
        _mm_storeu_si128 ( (__m128i*)&RGBA[i * 4], rgb32Pixel );
    }
}

方法2

void convertRGB24itoRGBA32i ( int width, int height, const RT_UByte *RGB, RT_UByte *RGBA )
{
    const size_t numPixels = (width - 1) * (height - 1);

    // Create the shuffle control mask for converting BGR to RGBA
    __m128i shuffleMask = _mm_setr_epi8 ( 2, 1, 0, 3, 5, 4, 3, 7, 8, 11, 10, 9, 13, 12, 15, 14 );

    for ( size_t i = 0; i < numPixels; i++ ) {
        __m128i sourcePixel = _mm_loadu_si128 ( reinterpret_cast<const __m128i*>(&RGB[i * 3]) );

        __m128i rgbaPixel = _mm_shuffle_epi8 ( sourcePixel, shuffleMask );

        __m128i alphaChannel = _mm_set1_epi32 ( 0xFF000000 );

        // Merge the RGBA channels
        rgbaPixel = _mm_or_si128 ( alphaChannel, rgbaPixel );

        _mm_storeu_si128 ( reinterpret_cast<__m128i*>(&RGBA[i * 4]), rgbaPixel );
    }
}

方法3

inline void convertBGRi24toBGRAi32 ( const ubyte3 *bgri24, ubyte4* bgrai32, t_size size )
{
    for ( ; size != 0; --size, ++bgrai32, ++bgri24 )
    {
        bgrai32->x = bgri24->x;
        bgrai32->y = bgri24->y;
        bgrai32->z = bgri24->z;
        bgrai32->w = 0xff;
    };

}

现寻求优化RGB转RGBA步骤的方法,以降低耗时。

内容的提问来源于stack exchange,提问作者bluetooth16

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 20:17:21