You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何空CUDA Kernel耗时远超CPU端OpenCV操作?CUDA图像处理价值何在?

问题背景

我了解到有时基于GPU的CUDA实现会比CPU端实现耗时更长,原因包括设备内存分配时间、数据往返内存的传输时间。为此我编写了两个脚本,仅测试CUDA Kernel的耗时(排除内存分配与传输环节),且使用的是一个空Kernel(无任何复杂运算)。但测试结果显示,该空Kernel的耗时仍是CPU端OpenCV操作的10倍。

测试代码

PyCUDA脚本

import cv2
import numpy as np
import pycuda.autoinit
import pycuda.driver as cuda
from pycuda.compiler import SourceModule
import argparse

parser = argparse.ArgumentParser()
parser.add_argument('--show', action='store_true',help="show the video while making")
parser.add_argument('--resize',type=int,default=800,help="if resize is needed")
parser.add_argument('--noconvert', action='store_true',help="avoid rgb conversion")

# Parse and print the results
args = parser.parse_args()
print(args)

# Path to the input H.264 file
input_video_path = '70secsmovie.h264'  # Replace with the path to your input H.264 file

# Path to the output MP4 file
output_video_path = 'output_video_cuda.mp4'  # Replace with the desired output MP4 file name

# Open the input video file
video_capture = cv2.VideoCapture(input_video_path)

# Check if the video file was opened successfully
if not video_capture.isOpened():
    print("Failed to open the video file.")
    exit()

# Set the desired width for the output frames
desired_width = args.resize #800

# Get the video properties
frame_width = int(video_capture.get(cv2.CAP_PROP_FRAME_WIDTH))
frame_height = int(video_capture.get(cv2.CAP_PROP_FRAME_HEIGHT))
fps = int(video_capture.get(cv2.CAP_PROP_FPS))

aspect_ratio = frame_width / frame_height
desired_height = int(desired_width / aspect_ratio)

# Create a VideoWriter object to save the output video
codec = cv2.VideoWriter_fourcc(*'mp4v')
# output_video = cv2.VideoWriter(output_video_path, codec, fps, (frame_width, frame_height))
output_video = cv2.VideoWriter(output_video_path, codec, fps, (desired_width, desired_height))


# Load the CUDA kernel for drawing the rectangle
mod = SourceModule("""
    __global__ void draw_rectangle_kernel(unsigned char *image, int image_width, int x, int y, int width, int height, unsigned char *color)
    {
        int row = blockIdx.y * blockDim.y + threadIdx.y;
        int col = blockIdx.x * blockDim.x + threadIdx.x;

        if (row >= y && row < y + height && col >= x && col < x + width)
        {
             // Perform no operation
        }
    }
""")

draw_rectangle_kernel = mod.get_function("draw_rectangle_kernel")

# Set the block dimensions
block_dim_x, block_dim_y = 16, 16

# Calculate the grid dimensions
grid_dim_x = (frame_width + block_dim_x - 1) // block_dim_x
grid_dim_y = (frame_height + block_dim_y - 1) // block_dim_y


# Define the rectangle properties (you can modify these as desired)
x, y, width, height = 100, 100, 200, 150
color = np.array([0, 255, 0], dtype=np.uint8)

# Initialize the frame count
frame_count = 0
average = 0.0 
start = cuda.Event()
end = cuda.Event()

# Read, process, and write each frame from the input video
while True:
    # Read a frame from the video file
    ret, frame = video_capture.read()

    # If the frame was not read successfully, the end of the video file is reached
    if not ret:
        break

    # Increment the frame count
    frame_count += 1


    if not args.noconvert:
        # Convert the frame to the RGB format
        frame_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
    else:
        frame_rgb = frame

    # start.record()
    # start.synchronize()

    # Upload the frame to the GPU
    frame_gpu = cuda.mem_alloc(frame_rgb.nbytes)
    cuda.memcpy_htod(frame_gpu, frame_rgb)

    start.record()
    start.synchronize()

    # Invoke the CUDA kernel to draw the rectangle
    # grid_dim_x = (frame_width + block_dim_x - 1) // block_dim_x
    # grid_dim_y = (frame_height + block_dim_y - 1) // block_dim_y
    draw_rectangle_kernel(frame_gpu, np.int32(frame_width), np.int32(x), np.int32(y),
                          np.int32(width), np.int32(height), cuda.In(color), block=(block_dim_x, block_dim_y, 1),
                          grid=(grid_dim_x, grid_dim_y))

    end.record()
    end.synchronize()

    # Download the modified frame from the GPU
    frame_with_rectangle_rgb = np.empty_like(frame_rgb)
    cuda.memcpy_dtoh(frame_with_rectangle_rgb, frame_gpu)

    # end.record()
    # end.synchronize()
    secs = start.time_till(end)*1e-3
    # print("Time of Squaring on GPU with inout")
    # print("%fs" % (secs))

    average = average + secs


    if not args.noconvert:
        # Convert the modified frame back to the BGR format
        frame_with_rectangle_bgr = cv2.cvtColor(frame_with_rectangle_rgb, cv2.COLOR_RGB2BGR)
    else:
        frame_with_rectangle_bgr = frame_with_rectangle_rgb



    # Resize the frame to the desired width and height while maintaining the aspect ratio
    resized_frame = cv2.resize(frame_with_rectangle_bgr, (desired_width, desired_height))


    # Write the modified frame to the output video
    output_video.write(resized_frame)


    # Write the modified frame to the output video
    # output_video.write(frame_with_rectangle_bgr)

    if args.show:
        # Display the modified frame (optional)
        # cv2.imshow('Modified Frame', frame_with_rectangle_bgr)
        cv2.imshow('Modified Frame', resized_frame)
        # Wait for the 'q' key to be pressed to stop (optional)
        if cv2.waitKey(1) & 0xFF == ord('q'):
            break

# Release the video capture and writer objects and close any open windows
video_capture.release()
output_video.release()
if args.show:
    cv2.destroyAllWindows()

# Print the total frame count
print("Total frames processed:", frame_count)

print("Operation took ", (average/frame_count))

OpenCV对比脚本

import cv2
import argparse

parser = argparse.ArgumentParser()
parser.add_argument('--show', action='store_true',help="show the video while making")

# Parse and print the results
args = parser.parse_args()
print(args)

# Path to the input H.264 file
input_video_path = '70secsmovie.h264'  # Replace with the path to your input H.264 file

# Path to the output MP4 file
output_video_path = 'output_video.mp4'  # Replace with the desired output MP4 file name

# Open the input video file
video_capture = cv2.VideoCapture(input_video_path)

# Check if the video file was opened successfully
if not video_capture.isOpened():
    print("Failed to open the video file.")
    exit()

# Set the desired width for the output frames
desired_width = 800

# Get the video properties
frame_width = int(video_capture.get(cv2.CAP_PROP_FRAME_WIDTH))
frame_height = int(video_capture.get(cv2.CAP_PROP_FRAME_HEIGHT))
fps = int(video_capture.get(cv2.CAP_PROP_FPS))
codec = cv2.VideoWriter_fourcc(*'mp4v')
aspect_ratio = frame_width / frame_height
desired_height = int(desired_width / aspect_ratio)

# Create a VideoWriter object to save the output video
# output_video = cv2.VideoWriter(output_video_path, codec, fps, (frame_width, frame_height))
output_video = cv2.VideoWriter(output_video_path, codec, fps, (desired_width, desired_height))

# Initialize the frame count
frame_count = 0
average = 0.0

# Read, process, and write each frame from the input video
while True:
    # Read a frame from the video file
    ret, frame = video_capture.read()

    # If the frame was not read successfully, the end of the video file is reached
    if not ret:
        break

    # Increment the frame count
    frame_count += 1

    # Draw a rectangle on the frame (you can modify the rectangle's properties here)
    x, y, width, height = 100, 100, 200, 150

    start = cv2.getTickCount()

    cv2.rectangle(frame, (x, y), (x + width, y + height), (0, 255, 0), 2)

    end = cv2.getTickCount()
    time = (end - start)/ cv2.getTickFrequency()
    # print("Time for Drawing Rectangle using OpenCV")
    # print("%fs" % (time))

    average = average + time

    # Resize the frame to the desired width and height while maintaining the aspect ratio
    resized_frame = cv2.resize(frame, (desired_width, desired_height))


    # Write the modified frame to the output video
    output_video.write(resized_frame)


    if args.show:
        # Display the modified frame (optional)
        cv2.imshow('Modified Frame', resized_frame)

        # Wait for the 'q' key to be pressed to stop (optional)
        if cv2.waitKey(1) & 0xFF == ord('q'):
            break

# Release the video capture and writer objects and close any open windows
video_capture.release()
output_video.release()
if args.show:
    cv2.destroyAllWindows()

# Print the total frame count
print("Total frames processed:", frame_count)

print("Operation took ", (average/frame_count))

测试结果

OpenCV脚本结果

Total frames processed: 704
Operation took  3.119159232954547e-05

PyCUDA脚本结果

Total frames processed: 704
Operation took  0.0003763223639266063
回答

你的测试用例刚好踩中了CUDA的“劣势场景”——空Kernel或者极简单运算,这时候CUDA的启动开销就会完全盖过它的并行优势。CUDA在图像处理中的实用性,得看它擅长的场景:

  • 大规模并行的复杂运算:当你需要对图像中每个像素(或大量像素)执行相同的复杂计算时,CUDA的优势才会显现。比如深度学习推理的卷积操作、大尺寸高斯模糊、SIFT特征提取等,GPU的数千个核心可以同时处理,速度是CPU的几十上百倍。
  • 批量处理任务:如果是处理成百上千张图片,或是4K/8K等高分辨率视频,CUDA可以把数据一次性放到GPU显存里,避免频繁的内存传输开销,连续处理所有帧/图片,整体效率会远超CPU。你当前的测试每帧都单独分配显存,其实可以优化为把显存分配移到循环外,重复使用同一块GPU内存,减少额外开销。
  • 自定义复杂算法:OpenCV提供的都是通用算法,如果你有自己定制的图像处理逻辑(比如特定的图像分割、自定义滤波),CUDA可以让你把这些逻辑写成Kernel,利用GPU并行加速,而CPU上实现同样的逻辑可能会慢很多。
  • 高分辨率图像处理:对于2K、4K甚至更高分辨率的图像,单个图像的像素数量达数百万甚至上亿,CPU逐个处理像素的速度会非常慢,而GPU可以同时处理数千个像素,直接把处理时间从秒级压缩到毫秒级。

简单说,CUDA不是用来替代CPU做简单操作的,而是用来解决CPU搞不定或者搞起来太慢的大规模、高复杂度图像处理任务。你的空Kernel测试,本质上是在测CUDA的启动延迟,这本来就不是它的强项。


内容的提问来源于stack exchange,提问作者KansaiRobot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 16:18:15