Swift中用Accelerate库实现带步长2D卷积及下采样优化
Great question! Your 2x2 downsampling (extracting every even-indexed pixel in both rows and columns) is a common operation, and the good news is there are much more efficient ways to handle this with Accelerate than your current loop or the convolution approach you tried. Let’s break down the best options:
Option 1: Use Planar Conversion + vDSP_deci (Most Efficient)
Your input is in interleaved ARGBFFFF format (A,R,G,B packed per pixel), which can be split into planar channels (one float array per channel) easily with vImage. Once planar, we can leverage vDSP_deci—a highly optimized SIMD function built specifically for fixed-stride decimation (extracting every nth element):
- Convert interleaved to planar:
UsevImageConvert_ARGBFFFFToPlanarFto split your sourcevImageBufferinto four separate float arrays (one for each A/R/G/B channel). This operation is hardware-accelerated and has minimal overhead. - Decimate each planar channel:
- First, process each row of the planar channel with
vDSP_deci(decimation factor = 2) to cut the width in half (keeping only even columns). - Then, extract every 2nd row from the resulting data to reduce the height by half (keeping only even rows). Since planar rows are contiguous, this can be done with
vDSP_vcopyusing a source stride equal to twice the row length of the halfway-decimated channel.
- First, process each row of the planar channel with
- Convert back to interleaved:
UsevImageConvert_PlanarFToARGBFFFFto combine the decimated planar channels back into your targetvImageBufferof sizefloor((M+1)/2) × floor((N+1)/2).
This approach uses Accelerate’s most optimized functions for both channel conversion and decimation, so it’ll be drastically faster than manual loops or convolution workarounds.
Option 2: Direct Interleaved Processing with vDSP_vcopy
If you want to skip planar conversion, you can work directly with the interleaved ARGBFFFF data using vDSP_vcopy, which supports custom source/destination strides:
- For each target row
i, the corresponding source row isi × 2(we skip all odd source rows entirely). - Treat each row as an array of
Pixel_FFFFelements (each containing 4 floats). To extract every 2nd pixel (even columns), callvDSP_vcopyf(for float data) with:- Source stride = 2 (skip every other
Pixel_FFFFelement) - Destination stride = 1 (write elements consecutively)
- Count =
floor((N+1)/2)(number of pixels per target row)
- Source stride = 2 (skip every other
Since vDSP_vcopy is vectorized, it handles SIMD operations under the hood—even with the row loop, this will outperform your original scalar loop by a wide margin.
Why Your Previous Approaches Didn’t Work (and What You Missed)
vImageConvolve_ARGBFFFFdoesn’t support strides because convolution is designed to slide the kernel over every pixel position—strides aren’t part of its intended use case, so it’s a poor fit for simple decimation.vDSP_vgathrrequires building index vectors, which adds overhead that cancels out any performance gains.vDSP_deciandvDSP_vcopyare purpose-built for fixed-stride extraction and avoid this overhead entirely.- You don’t need 2D convolution here at all! Your operation is a simple decimation, not a filtered downsample—using decimation-specific functions is far more efficient than forcing a convolution solution.
Quick Code Snippet (Planar Approach Outline)
// Assume srcBuffer is your input vImageBuffer (ARGBFFFF), dstBuffer is pre-allocated target vImage_Buffer planarBuffers[4]; // Initialize planar buffers with width = srcBuffer.width, height = srcBuffer.height // (Allocate memory for each planar channel using malloc or vImage helpers) // Convert interleaved ARGB to planar float channels vImageConvert_ARGBFFFFToPlanarF(&srcBuffer, planarBuffers[0].data, planarBuffers[1].data, planarBuffers[2].data, planarBuffers[3].data, kvImageNoFlags); // Decimate width for each channel vImage_Buffer halfWidthBuffers[4]; // Initialize with width = floor((srcBuffer.width +1)/2), height = srcBuffer.height for (int c = 0; c < 4; c++) { vDSP_deci(planarBuffers[c].data, 2, halfWidthBuffers[c].data, 1, halfWidthBuffers[c].width * halfWidthBuffers[c].height, 2); } // Decimate height (extract every 2nd row) vImage_Buffer finalBuffers[4]; // Initialize with width = halfWidthBuffers[0].width, height = floor((srcBuffer.height +1)/2) for (int c = 0; c <4; c++) { vDSP_vcopy(halfWidthBuffers[c].data, halfWidthBuffers[c].width * 2, finalBuffers[c].data, finalBuffers[c].width, finalBuffers[c].width * finalBuffers[c].height); } // Convert back to interleaved ARGBFFFF vImageConvert_PlanarFToARGBFFFF(finalBuffers[0].data, finalBuffers[1].data, finalBuffers[2].data, finalBuffers[3].data, &dstBuffer, kvImageNoFlags); // Don't forget to free allocated planar buffers!
内容的提问来源于stack exchange,提问作者Keen R.D.

