基于Numpy实现边界框内容快速复制到画布的优化方法
优化多边界框组合图像生成的性能
问题背景
需要从大尺寸图像中提取不同边界框的组合,生成仅包含选中框区域(其余为白色)的图像并保存。当前核心耗时在于每次循环复制全尺寸空白画布,再逐个复制框区域,单循环耗时0.227秒,需执行数千次,性能瓶颈明显。
优化方案
1. 只操作最终需要的裁剪区域,避免全画布复制
原流程先复制全尺寸空白画布,复制所有框后再裁剪,会浪费大量时间在无关区域的处理上。可以先计算当前框组合的最小包围区域,直接在这个小尺寸画布上操作:
import numpy as np import cv2 import time orig = np.zeros((9536, 13480, 3), dtype=np.uint8) rects = [[1000,1000,1100,1100], [1100,1000,1200,1100],[1100,1100,1200,1200]] loops = 10 # 预定义需要测试的框组合(1+2、1+3、2+3) combinations = [ [rects[0], rects[1]], [rects[0], rects[2]], [rects[1], rects[2]] ] starttime = time.perf_counter() for i in range(loops): # 轮询框组合 current_rects = combinations[i % len(combinations)] # 计算当前组合的最小包围盒 xs = np.array([r[0] for r in current_rects] + [r[2] for r in current_rects]) ys = np.array([r[1] for r in current_rects] + [r[3] for r in current_rects]) min_x, max_x = xs.min(), xs.max() min_y, max_y = ys.min(), ys.max() # 生成对应大小的白色画布 canvas = np.full((max_y - min_y, max_x - min_x, 3), 255, dtype=np.uint8) # 复制框内容到画布(计算相对于包围盒的偏移) for rect in current_rects: x1, y1, x2, y2 = rect cx1 = x1 - min_x cy1 = y1 - min_y canvas[cy1:y2-min_y, cx1:x2-min_x] = orig[y1:y2, x1:x2] cv2.imwrite(f"{i}.jpg", canvas) fulltime = time.perf_counter()-starttime looptime = fulltime/loops print("Time taken per loop:: ", looptime)
优化逻辑:仅处理最终要保存的小区域,避免全尺寸画布的复制和无效区域操作,大幅降低内存操作量。
2. 预缓存所有框的图像切片,避免重复读取原图像
提前把每个边界框对应的图像区域提取出来缓存,循环时直接使用缓存的切片,减少重复的大数组切片开销:
import numpy as np import cv2 import time orig = np.zeros((9536, 13480, 3), dtype=np.uint8) rects = [[1000,1000,1100,1100], [1100,1000,1200,1100],[1100,1100,1200,1200]] loops = 10 # 预缓存每个框的图像区域 cached_slices = [] for x1,y1,x2,y2 in rects: cached_slices.append(orig[y1:y2, x1:x2].copy()) # 用索引标记框组合 combinations = [[0,1], [0,2], [1,2]] starttime = time.perf_counter() for i in range(loops): combo_idx = combinations[i % len(combinations)] current_rects = [rects[idx] for idx in combo_idx] current_slices = [cached_slices[idx] for idx in combo_idx] # 计算包围盒 xs = np.array([r[0] for r in current_rects] + [r[2] for r in current_rects]) ys = np.array([r[1] for r in current_rects] + [r[3] for r in current_rects]) min_x, max_x = xs.min(), xs.max() min_y, max_y = ys.min(), ys.max() canvas = np.full((max_y-min_y, max_x-min_x, 3), 255, dtype=np.uint8) # 用缓存切片复制内容 for rect, slice_img in zip(current_rects, current_slices): x1,y1,x2,y2 = rect cx1 = x1 - min_x cy1 = y1 - min_y canvas[cy1:cy1+slice_img.shape[0], cx1:cx1+slice_img.shape[1]] = slice_img cv2.imwrite(f"{i}.jpg", canvas) fulltime = time.perf_counter()-starttime looptime = fulltime/loops print("Time taken per loop:: ", looptime)
优化逻辑:原图像的切片操作属于高开销的内存IO,预缓存后每次循环仅需复制小尺寸预存切片,减少重复计算。
3. 使用掩码批量生成目标区域(适合多组合场景)
如果框组合数量极多,可以先创建掩码标记需要保留的区域,利用NumPy向量化操作批量生成目标图像,替代逐个框的循环复制:
import numpy as np import cv2 import time orig = np.zeros((9536, 13480, 3), dtype=np.uint8) rects = [[1000,1000,1100,1100], [1100,1000,1200,1100],[1100,1100,1200,1200]] loops = 10 combinations = [[0,1], [0,2], [1,2]] starttime = time.perf_counter() for i in range(loops): combo_idx = combinations[i % len(combinations)] current_rects = [rects[idx] for idx in combo_idx] # 计算包围盒 xs = np.array([r[0] for r in current_rects] + [r[2] for r in current_rects]) ys = np.array([r[1] for r in current_rects] + [r[3] for r in current_rects]) min_x, max_x = xs.min(), xs.max() min_y, max_y = ys.min(), ys.max() # 创建包围盒内的掩码 mask = np.zeros((max_y-min_y, max_x-min_x), dtype=bool) for rect in current_rects: x1,y1,x2,y2 = rect cx1 = x1 - min_x cy1 = y1 - min_y mask[cy1:y2-min_y, cx1:x2-min_x] = True # 批量生成目标画布 canvas = np.full((max_y-min_y, max_x-min_x, 3), 255, dtype=np.uint8) orig_crop = orig[min_y:max_y, min_x:max_x] canvas[mask] = orig_crop[mask] cv2.imwrite(f"{i}.jpg", canvas) fulltime = time.perf_counter()-starttime looptime = fulltime/loops print("Time taken per loop:: ", looptime)
优化逻辑:NumPy向量化操作的效率远高于Python循环,通过掩码批量赋值,大幅减少循环次数。
内容的提问来源于stack exchange,提问作者Jkind9
相关产品推荐
相关产品推荐

