如何通过cudaMemcpy将OpenCV HWC/NHWC格式图像转为CHW/NCHW格式拷贝至CUDA内存,且避免高开销的cv::split操作?
Absolutely! You can skip the costly cv::split operation entirely and directly convert your HWC-formatted OpenCV cv::Mat to CHW/NCHW format in CUDA memory—here are two efficient approaches tailored to different use cases:
1. Use cudaMemcpy2D with Custom Strides (Single Image HWC → CHW)
This method leverages CUDA's 2D memory copy function to directly extract each channel from the host-side HWC Mat to CUDA memory, without splitting channels on the host first.
Code Example
#include <opencv2/opencv.hpp> #include <cuda_runtime.h> int main() { // Load HWC-format OpenCV image (BGR channel order) cv::Mat usual_image = cv::imread("input.jpg", cv::IMREAD_COLOR); const int H = usual_image.rows; const int W = usual_image.cols; const int C = usual_image.channels(); // Allocate CUDA memory for CHW format (C × H × W) unsigned char* d_chw_image; const size_t chw_size = C * H * W * sizeof(unsigned char); cudaMalloc(&d_chw_image, chw_size); // Copy each channel individually using custom pitch/stride for (int c = 0; c < C; ++c) { // Source: Start at the c-th channel of the first pixel, stride past other channels per row const void* src_ptr = usual_image.data + c; const size_t src_pitch = W * C * sizeof(unsigned char); // Destination: Start at the beginning of the c-th channel block, stride per row of the single channel void* dst_ptr = d_chw_image + c * H * W; const size_t dst_pitch = W * sizeof(unsigned char); // Copy entire channel in one 2D transfer cudaMemcpy2D(dst_ptr, dst_pitch, src_ptr, src_pitch, W * sizeof(unsigned char), H, cudaMemcpyHostToDevice); } // ...后续CUDA操作... // Cleanup cudaFree(d_chw_image); return 0; }
Why This Works
Instead of splitting the image into separate host-side Mats (like cv::split does), we use the inherent stride of the HWC format to skip non-target channel data during the copy. This eliminates extra host memory usage and the overhead of multiple host-side memory allocations/copies.
2. Custom CUDA Kernel (Batch NHWC → NCHW or Large Images)
For batch processing (NHWC) or maximum efficiency with large images, a lightweight CUDA kernel can handle the format conversion entirely on the GPU, after a single host-to-device copy of the raw HWC data.
Code Example
// CUDA Kernel: Convert HWC to CHW in-place on device __global__ void hwc_to_chw_kernel(const unsigned char* d_hwc, unsigned char* d_chw, int H, int W, int C) { // Calculate pixel coordinates for current thread const int h = blockIdx.y * blockDim.y + threadIdx.y; const int w = blockIdx.x * blockDim.x + threadIdx.x; if (h >= H || w >= W) return; // Map HWC (h,w,c) index to CHW (c,h,w) index for (int c = 0; c < C; ++c) { const int hwc_idx = h * W * C + w * C + c; const int chw_idx = c * H * W + h * W + w; d_chw[chw_idx] = d_hwc[hwc_idx]; } } // Host-side invocation int main() { cv::Mat usual_image = cv::imread("input.jpg", cv::IMREAD_COLOR); const int H = usual_image.rows; const int W = usual_image.cols; const int C = usual_image.channels(); // Copy raw HWC data to CUDA memory (for NHWC, just append more images sequentially) unsigned char* d_hwc_image; const size_t hwc_size = H * W * C * sizeof(unsigned char); cudaMalloc(&d_hwc_image, hwc_size); cudaMemcpy(d_hwc_image, usual_image.data, hwc_size, cudaMemcpyHostToDevice); // Allocate CHW-format CUDA memory unsigned char* d_chw_image; const size_t chw_size = C * H * W * sizeof(unsigned char); cudaMalloc(&d_chw_image, chw_size); // Launch kernel: 32x32 thread blocks (standard for GPU efficiency) const dim3 block_size(32, 32); const dim3 grid_size((W + block_size.x - 1) / block_size.x, (H + block_size.y - 1) / block_size.y); hwc_to_chw_kernel<<<grid_size, block_size>>>(d_hwc_image, d_chw_image, H, W, C); cudaDeviceSynchronize(); // Optional: Wait for kernel completion (for debugging) // ...后续CUDA操作... // Cleanup cudaFree(d_hwc_image); cudaFree(d_chw_image); return 0; }
Key Advantages Over cv::split
- No extra host memory bloat:
cv::splitcreates C separate host-side Mats, doubling your host memory usage. These methods avoid that entirely. - Faster execution: Both approaches minimize host-side processing. The kernel method leverages GPU parallelism for batch or large-image workloads.
- Flexibility: The kernel can easily be modified to handle NHWC→NCHW by adding a batch dimension to the index calculations.
内容的提问来源于stack exchange,提问作者SofaScience

