为何NumPy实现BGR转灰度比OpenCV cvtColor慢18倍?
为什么OpenCV的BGR转灰度比NumPy实现快18倍?
问题背景
根据OpenCV文档,它通过线性变换 Y = 0.299R + 0.587G + 0.114B 将BGR图像转为灰度图。我尝试用NumPy模拟该过程:将HxWx3的BGR矩阵与3x1的系数向量[0.114, 0.587, 0.299]'相乘,得到HxWx1的灰度矩阵。
对应的NumPy代码如下:
import cv2 import numpy as np import time im = cv2.imread(IM_PATHS[0], cv2.IMREAD_COLOR) # 预分配灰度图内存 dst = np.zeros(im.shape[:2], dtype = np.uint8) # BGR转灰度的列向量系数 bgr_weight_arr = np.array((0.114,0.587,0.299), dtype = np.float32).reshape(3,1) for im_path in IM_PATHS: im = cv2.imread(im_path , cv2.IMREAD_COLOR) t1 = time.time() # NumPy矩阵乘法实现转换 dst[:,:] = (im @ bgr_weight_arr).reshape(*dst.shape) t2 = time.time() print(f'runtime: {(t2-t1):.3f}sec')
处理12MP(4000x3000像素)图像时,上述NumPy流程每张图耗时约90ms(未对结果取整)。而替换为OpenCV的dst[:,:] = cv2.cvtColor(im, cv2.COLOR_BGR2GRAY)后,每张图仅需约5ms,速度快18倍!
我一直认为NumPy会利用SIMD等加速技术,为何OpenCV能快这么多?
更新:尝试量化乘法后的结果
即便使用量化乘法,NumPy的耗时仍维持在90ms左右,代码如下:
rgb_weight_arr_uint16 = np.round(256 * np.array((0.114,0.587,0.299))).astype('uint16').reshape(3,1) for im_path in IM_PATHS: im = cv2.imread(im_path , cv2.IMREAD_COLOR) t1 = time.time() # NumPy量化乘法实现转换 dst[:,:] = np.right_shift(im @ bgr_weight_arr_uint16, 8).reshape(*dst.shape) t2 = time.time() print(f'runtime: {(t2-t1):.3f}sec')
内容的提问来源于stack exchange,提问作者SomethingSomething
相关产品推荐
相关产品推荐

