为何空CUDA Kernel耗时远超CPU端OpenCV操作?CUDA图像处理价值何在?
问题背景
我了解到有时基于GPU的CUDA实现会比CPU端实现耗时更长,原因包括设备内存分配时间、数据往返内存的传输时间。为此我编写了两个脚本,仅测试CUDA Kernel的耗时(排除内存分配与传输环节),且使用的是一个空Kernel(无任何复杂运算)。但测试结果显示,该空Kernel的耗时仍是CPU端OpenCV操作的10倍。
测试代码
PyCUDA脚本
import cv2 import numpy as np import pycuda.autoinit import pycuda.driver as cuda from pycuda.compiler import SourceModule import argparse parser = argparse.ArgumentParser() parser.add_argument('--show', action='store_true',help="show the video while making") parser.add_argument('--resize',type=int,default=800,help="if resize is needed") parser.add_argument('--noconvert', action='store_true',help="avoid rgb conversion") # Parse and print the results args = parser.parse_args() print(args) # Path to the input H.264 file input_video_path = '70secsmovie.h264' # Replace with the path to your input H.264 file # Path to the output MP4 file output_video_path = 'output_video_cuda.mp4' # Replace with the desired output MP4 file name # Open the input video file video_capture = cv2.VideoCapture(input_video_path) # Check if the video file was opened successfully if not video_capture.isOpened(): print("Failed to open the video file.") exit() # Set the desired width for the output frames desired_width = args.resize #800 # Get the video properties frame_width = int(video_capture.get(cv2.CAP_PROP_FRAME_WIDTH)) frame_height = int(video_capture.get(cv2.CAP_PROP_FRAME_HEIGHT)) fps = int(video_capture.get(cv2.CAP_PROP_FPS)) aspect_ratio = frame_width / frame_height desired_height = int(desired_width / aspect_ratio) # Create a VideoWriter object to save the output video codec = cv2.VideoWriter_fourcc(*'mp4v') # output_video = cv2.VideoWriter(output_video_path, codec, fps, (frame_width, frame_height)) output_video = cv2.VideoWriter(output_video_path, codec, fps, (desired_width, desired_height)) # Load the CUDA kernel for drawing the rectangle mod = SourceModule(""" __global__ void draw_rectangle_kernel(unsigned char *image, int image_width, int x, int y, int width, int height, unsigned char *color) { int row = blockIdx.y * blockDim.y + threadIdx.y; int col = blockIdx.x * blockDim.x + threadIdx.x; if (row >= y && row < y + height && col >= x && col < x + width) { // Perform no operation } } """) draw_rectangle_kernel = mod.get_function("draw_rectangle_kernel") # Set the block dimensions block_dim_x, block_dim_y = 16, 16 # Calculate the grid dimensions grid_dim_x = (frame_width + block_dim_x - 1) // block_dim_x grid_dim_y = (frame_height + block_dim_y - 1) // block_dim_y # Define the rectangle properties (you can modify these as desired) x, y, width, height = 100, 100, 200, 150 color = np.array([0, 255, 0], dtype=np.uint8) # Initialize the frame count frame_count = 0 average = 0.0 start = cuda.Event() end = cuda.Event() # Read, process, and write each frame from the input video while True: # Read a frame from the video file ret, frame = video_capture.read() # If the frame was not read successfully, the end of the video file is reached if not ret: break # Increment the frame count frame_count += 1 if not args.noconvert: # Convert the frame to the RGB format frame_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB) else: frame_rgb = frame # start.record() # start.synchronize() # Upload the frame to the GPU frame_gpu = cuda.mem_alloc(frame_rgb.nbytes) cuda.memcpy_htod(frame_gpu, frame_rgb) start.record() start.synchronize() # Invoke the CUDA kernel to draw the rectangle # grid_dim_x = (frame_width + block_dim_x - 1) // block_dim_x # grid_dim_y = (frame_height + block_dim_y - 1) // block_dim_y draw_rectangle_kernel(frame_gpu, np.int32(frame_width), np.int32(x), np.int32(y), np.int32(width), np.int32(height), cuda.In(color), block=(block_dim_x, block_dim_y, 1), grid=(grid_dim_x, grid_dim_y)) end.record() end.synchronize() # Download the modified frame from the GPU frame_with_rectangle_rgb = np.empty_like(frame_rgb) cuda.memcpy_dtoh(frame_with_rectangle_rgb, frame_gpu) # end.record() # end.synchronize() secs = start.time_till(end)*1e-3 # print("Time of Squaring on GPU with inout") # print("%fs" % (secs)) average = average + secs if not args.noconvert: # Convert the modified frame back to the BGR format frame_with_rectangle_bgr = cv2.cvtColor(frame_with_rectangle_rgb, cv2.COLOR_RGB2BGR) else: frame_with_rectangle_bgr = frame_with_rectangle_rgb # Resize the frame to the desired width and height while maintaining the aspect ratio resized_frame = cv2.resize(frame_with_rectangle_bgr, (desired_width, desired_height)) # Write the modified frame to the output video output_video.write(resized_frame) # Write the modified frame to the output video # output_video.write(frame_with_rectangle_bgr) if args.show: # Display the modified frame (optional) # cv2.imshow('Modified Frame', frame_with_rectangle_bgr) cv2.imshow('Modified Frame', resized_frame) # Wait for the 'q' key to be pressed to stop (optional) if cv2.waitKey(1) & 0xFF == ord('q'): break # Release the video capture and writer objects and close any open windows video_capture.release() output_video.release() if args.show: cv2.destroyAllWindows() # Print the total frame count print("Total frames processed:", frame_count) print("Operation took ", (average/frame_count))
OpenCV对比脚本
import cv2 import argparse parser = argparse.ArgumentParser() parser.add_argument('--show', action='store_true',help="show the video while making") # Parse and print the results args = parser.parse_args() print(args) # Path to the input H.264 file input_video_path = '70secsmovie.h264' # Replace with the path to your input H.264 file # Path to the output MP4 file output_video_path = 'output_video.mp4' # Replace with the desired output MP4 file name # Open the input video file video_capture = cv2.VideoCapture(input_video_path) # Check if the video file was opened successfully if not video_capture.isOpened(): print("Failed to open the video file.") exit() # Set the desired width for the output frames desired_width = 800 # Get the video properties frame_width = int(video_capture.get(cv2.CAP_PROP_FRAME_WIDTH)) frame_height = int(video_capture.get(cv2.CAP_PROP_FRAME_HEIGHT)) fps = int(video_capture.get(cv2.CAP_PROP_FPS)) codec = cv2.VideoWriter_fourcc(*'mp4v') aspect_ratio = frame_width / frame_height desired_height = int(desired_width / aspect_ratio) # Create a VideoWriter object to save the output video # output_video = cv2.VideoWriter(output_video_path, codec, fps, (frame_width, frame_height)) output_video = cv2.VideoWriter(output_video_path, codec, fps, (desired_width, desired_height)) # Initialize the frame count frame_count = 0 average = 0.0 # Read, process, and write each frame from the input video while True: # Read a frame from the video file ret, frame = video_capture.read() # If the frame was not read successfully, the end of the video file is reached if not ret: break # Increment the frame count frame_count += 1 # Draw a rectangle on the frame (you can modify the rectangle's properties here) x, y, width, height = 100, 100, 200, 150 start = cv2.getTickCount() cv2.rectangle(frame, (x, y), (x + width, y + height), (0, 255, 0), 2) end = cv2.getTickCount() time = (end - start)/ cv2.getTickFrequency() # print("Time for Drawing Rectangle using OpenCV") # print("%fs" % (time)) average = average + time # Resize the frame to the desired width and height while maintaining the aspect ratio resized_frame = cv2.resize(frame, (desired_width, desired_height)) # Write the modified frame to the output video output_video.write(resized_frame) if args.show: # Display the modified frame (optional) cv2.imshow('Modified Frame', resized_frame) # Wait for the 'q' key to be pressed to stop (optional) if cv2.waitKey(1) & 0xFF == ord('q'): break # Release the video capture and writer objects and close any open windows video_capture.release() output_video.release() if args.show: cv2.destroyAllWindows() # Print the total frame count print("Total frames processed:", frame_count) print("Operation took ", (average/frame_count))
测试结果
OpenCV脚本结果
Total frames processed: 704 Operation took 3.119159232954547e-05
PyCUDA脚本结果
Total frames processed: 704 Operation took 0.0003763223639266063
回答
你的测试用例刚好踩中了CUDA的“劣势场景”——空Kernel或者极简单运算,这时候CUDA的启动开销就会完全盖过它的并行优势。CUDA在图像处理中的实用性,得看它擅长的场景:
- 大规模并行的复杂运算:当你需要对图像中每个像素(或大量像素)执行相同的复杂计算时,CUDA的优势才会显现。比如深度学习推理的卷积操作、大尺寸高斯模糊、SIFT特征提取等,GPU的数千个核心可以同时处理,速度是CPU的几十上百倍。
- 批量处理任务:如果是处理成百上千张图片,或是4K/8K等高分辨率视频,CUDA可以把数据一次性放到GPU显存里,避免频繁的内存传输开销,连续处理所有帧/图片,整体效率会远超CPU。你当前的测试每帧都单独分配显存,其实可以优化为把显存分配移到循环外,重复使用同一块GPU内存,减少额外开销。
- 自定义复杂算法:OpenCV提供的都是通用算法,如果你有自己定制的图像处理逻辑(比如特定的图像分割、自定义滤波),CUDA可以让你把这些逻辑写成Kernel,利用GPU并行加速,而CPU上实现同样的逻辑可能会慢很多。
- 高分辨率图像处理:对于2K、4K甚至更高分辨率的图像,单个图像的像素数量达数百万甚至上亿,CPU逐个处理像素的速度会非常慢,而GPU可以同时处理数千个像素,直接把处理时间从秒级压缩到毫秒级。
简单说,CUDA不是用来替代CPU做简单操作的,而是用来解决CPU搞不定或者搞起来太慢的大规模、高复杂度图像处理任务。你的空Kernel测试,本质上是在测CUDA的启动延迟,这本来就不是它的强项。
内容的提问来源于stack exchange,提问作者KansaiRobot
相关产品推荐
相关产品推荐

