You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过cudaMemcpy将OpenCV HWC/NHWC格式图像转为CHW/NCHW格式拷贝至CUDA内存,且避免高开销的cv::split操作?

Absolutely! You can skip the costly cv::split operation entirely and directly convert your HWC-formatted OpenCV cv::Mat to CHW/NCHW format in CUDA memory—here are two efficient approaches tailored to different use cases:

1. Use cudaMemcpy2D with Custom Strides (Single Image HWC → CHW)

This method leverages CUDA's 2D memory copy function to directly extract each channel from the host-side HWC Mat to CUDA memory, without splitting channels on the host first.

Code Example

#include <opencv2/opencv.hpp>
#include <cuda_runtime.h>

int main() {
    // Load HWC-format OpenCV image (BGR channel order)
    cv::Mat usual_image = cv::imread("input.jpg", cv::IMREAD_COLOR);
    const int H = usual_image.rows;
    const int W = usual_image.cols;
    const int C = usual_image.channels();

    // Allocate CUDA memory for CHW format (C × H × W)
    unsigned char* d_chw_image;
    const size_t chw_size = C * H * W * sizeof(unsigned char);
    cudaMalloc(&d_chw_image, chw_size);

    // Copy each channel individually using custom pitch/stride
    for (int c = 0; c < C; ++c) {
        // Source: Start at the c-th channel of the first pixel, stride past other channels per row
        const void* src_ptr = usual_image.data + c;
        const size_t src_pitch = W * C * sizeof(unsigned char);

        // Destination: Start at the beginning of the c-th channel block, stride per row of the single channel
        void* dst_ptr = d_chw_image + c * H * W;
        const size_t dst_pitch = W * sizeof(unsigned char);

        // Copy entire channel in one 2D transfer
        cudaMemcpy2D(dst_ptr, dst_pitch,
                     src_ptr, src_pitch,
                     W * sizeof(unsigned char), H,
                     cudaMemcpyHostToDevice);
    }

    // ...后续CUDA操作...

    // Cleanup
    cudaFree(d_chw_image);
    return 0;
}

Why This Works

Instead of splitting the image into separate host-side Mats (like cv::split does), we use the inherent stride of the HWC format to skip non-target channel data during the copy. This eliminates extra host memory usage and the overhead of multiple host-side memory allocations/copies.

2. Custom CUDA Kernel (Batch NHWC → NCHW or Large Images)

For batch processing (NHWC) or maximum efficiency with large images, a lightweight CUDA kernel can handle the format conversion entirely on the GPU, after a single host-to-device copy of the raw HWC data.

Code Example

// CUDA Kernel: Convert HWC to CHW in-place on device
__global__ void hwc_to_chw_kernel(const unsigned char* d_hwc, unsigned char* d_chw, int H, int W, int C) {
    // Calculate pixel coordinates for current thread
    const int h = blockIdx.y * blockDim.y + threadIdx.y;
    const int w = blockIdx.x * blockDim.x + threadIdx.x;

    if (h >= H || w >= W) return;

    // Map HWC (h,w,c) index to CHW (c,h,w) index
    for (int c = 0; c < C; ++c) {
        const int hwc_idx = h * W * C + w * C + c;
        const int chw_idx = c * H * W + h * W + w;
        d_chw[chw_idx] = d_hwc[hwc_idx];
    }
}

// Host-side invocation
int main() {
    cv::Mat usual_image = cv::imread("input.jpg", cv::IMREAD_COLOR);
    const int H = usual_image.rows;
    const int W = usual_image.cols;
    const int C = usual_image.channels();

    // Copy raw HWC data to CUDA memory (for NHWC, just append more images sequentially)
    unsigned char* d_hwc_image;
    const size_t hwc_size = H * W * C * sizeof(unsigned char);
    cudaMalloc(&d_hwc_image, hwc_size);
    cudaMemcpy(d_hwc_image, usual_image.data, hwc_size, cudaMemcpyHostToDevice);

    // Allocate CHW-format CUDA memory
    unsigned char* d_chw_image;
    const size_t chw_size = C * H * W * sizeof(unsigned char);
    cudaMalloc(&d_chw_image, chw_size);

    // Launch kernel: 32x32 thread blocks (standard for GPU efficiency)
    const dim3 block_size(32, 32);
    const dim3 grid_size((W + block_size.x - 1) / block_size.x, 
                        (H + block_size.y - 1) / block_size.y);
    hwc_to_chw_kernel<<<grid_size, block_size>>>(d_hwc_image, d_chw_image, H, W, C);
    cudaDeviceSynchronize(); // Optional: Wait for kernel completion (for debugging)

    // ...后续CUDA操作...

    // Cleanup
    cudaFree(d_hwc_image);
    cudaFree(d_chw_image);
    return 0;
}

Key Advantages Over cv::split

  • No extra host memory bloat: cv::split creates C separate host-side Mats, doubling your host memory usage. These methods avoid that entirely.
  • Faster execution: Both approaches minimize host-side processing. The kernel method leverages GPU parallelism for batch or large-image workloads.
  • Flexibility: The kernel can easily be modified to handle NHWC→NCHW by adding a batch dimension to the index calculations.

内容的提问来源于stack exchange,提问作者SofaScience

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:55:21