You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Numpy实现边界框内容快速复制到画布的优化方法

优化多边界框组合图像生成的性能

问题背景

需要从大尺寸图像中提取不同边界框的组合,生成仅包含选中框区域(其余为白色)的图像并保存。当前核心耗时在于每次循环复制全尺寸空白画布,再逐个复制框区域,单循环耗时0.227秒,需执行数千次,性能瓶颈明显。

优化方案


1. 只操作最终需要的裁剪区域,避免全画布复制

原流程先复制全尺寸空白画布,复制所有框后再裁剪,会浪费大量时间在无关区域的处理上。可以先计算当前框组合的最小包围区域,直接在这个小尺寸画布上操作:

import numpy as np
import cv2
import time

orig = np.zeros((9536, 13480, 3), dtype=np.uint8)
rects = [[1000,1000,1100,1100], [1100,1000,1200,1100],[1100,1100,1200,1200]]
loops = 10

# 预定义需要测试的框组合(1+2、1+3、2+3)
combinations = [
    [rects[0], rects[1]],
    [rects[0], rects[2]],
    [rects[1], rects[2]]
]

starttime = time.perf_counter()

for i in range(loops):
    # 轮询框组合
    current_rects = combinations[i % len(combinations)]
    # 计算当前组合的最小包围盒
    xs = np.array([r[0] for r in current_rects] + [r[2] for r in current_rects])
    ys = np.array([r[1] for r in current_rects] + [r[3] for r in current_rects])
    min_x, max_x = xs.min(), xs.max()
    min_y, max_y = ys.min(), ys.max()
    # 生成对应大小的白色画布
    canvas = np.full((max_y - min_y, max_x - min_x, 3), 255, dtype=np.uint8)
    
    # 复制框内容到画布(计算相对于包围盒的偏移)
    for rect in current_rects:
        x1, y1, x2, y2 = rect
        cx1 = x1 - min_x
        cy1 = y1 - min_y
        canvas[cy1:y2-min_y, cx1:x2-min_x] = orig[y1:y2, x1:x2]
    
    cv2.imwrite(f"{i}.jpg", canvas)

fulltime = time.perf_counter()-starttime
looptime = fulltime/loops
print("Time taken per loop:: ", looptime)

优化逻辑:仅处理最终要保存的小区域,避免全尺寸画布的复制和无效区域操作,大幅降低内存操作量。


2. 预缓存所有框的图像切片,避免重复读取原图像

提前把每个边界框对应的图像区域提取出来缓存,循环时直接使用缓存的切片,减少重复的大数组切片开销:

import numpy as np
import cv2
import time

orig = np.zeros((9536, 13480, 3), dtype=np.uint8)
rects = [[1000,1000,1100,1100], [1100,1000,1200,1100],[1100,1100,1200,1200]]
loops = 10

# 预缓存每个框的图像区域
cached_slices = []
for x1,y1,x2,y2 in rects:
    cached_slices.append(orig[y1:y2, x1:x2].copy())

# 用索引标记框组合
combinations = [[0,1], [0,2], [1,2]]

starttime = time.perf_counter()

for i in range(loops):
    combo_idx = combinations[i % len(combinations)]
    current_rects = [rects[idx] for idx in combo_idx]
    current_slices = [cached_slices[idx] for idx in combo_idx]
    
    # 计算包围盒
    xs = np.array([r[0] for r in current_rects] + [r[2] for r in current_rects])
    ys = np.array([r[1] for r in current_rects] + [r[3] for r in current_rects])
    min_x, max_x = xs.min(), xs.max()
    min_y, max_y = ys.min(), ys.max()
    canvas = np.full((max_y-min_y, max_x-min_x, 3), 255, dtype=np.uint8)
    
    # 用缓存切片复制内容
    for rect, slice_img in zip(current_rects, current_slices):
        x1,y1,x2,y2 = rect
        cx1 = x1 - min_x
        cy1 = y1 - min_y
        canvas[cy1:cy1+slice_img.shape[0], cx1:cx1+slice_img.shape[1]] = slice_img
    
    cv2.imwrite(f"{i}.jpg", canvas)

fulltime = time.perf_counter()-starttime
looptime = fulltime/loops
print("Time taken per loop:: ", looptime)

优化逻辑:原图像的切片操作属于高开销的内存IO,预缓存后每次循环仅需复制小尺寸预存切片,减少重复计算。


3. 使用掩码批量生成目标区域(适合多组合场景)

如果框组合数量极多,可以先创建掩码标记需要保留的区域,利用NumPy向量化操作批量生成目标图像,替代逐个框的循环复制:

import numpy as np
import cv2
import time

orig = np.zeros((9536, 13480, 3), dtype=np.uint8)
rects = [[1000,1000,1100,1100], [1100,1000,1200,1100],[1100,1100,1200,1200]]
loops = 10

combinations = [[0,1], [0,2], [1,2]]

starttime = time.perf_counter()

for i in range(loops):
    combo_idx = combinations[i % len(combinations)]
    current_rects = [rects[idx] for idx in combo_idx]
    
    # 计算包围盒
    xs = np.array([r[0] for r in current_rects] + [r[2] for r in current_rects])
    ys = np.array([r[1] for r in current_rects] + [r[3] for r in current_rects])
    min_x, max_x = xs.min(), xs.max()
    min_y, max_y = ys.min(), ys.max()
    
    # 创建包围盒内的掩码
    mask = np.zeros((max_y-min_y, max_x-min_x), dtype=bool)
    for rect in current_rects:
        x1,y1,x2,y2 = rect
        cx1 = x1 - min_x
        cy1 = y1 - min_y
        mask[cy1:y2-min_y, cx1:x2-min_x] = True
    
    # 批量生成目标画布
    canvas = np.full((max_y-min_y, max_x-min_x, 3), 255, dtype=np.uint8)
    orig_crop = orig[min_y:max_y, min_x:max_x]
    canvas[mask] = orig_crop[mask]
    
    cv2.imwrite(f"{i}.jpg", canvas)

fulltime = time.perf_counter()-starttime
looptime = fulltime/loops
print("Time taken per loop:: ", looptime)

优化逻辑:NumPy向量化操作的效率远高于Python循环,通过掩码批量赋值,大幅减少循环次数。


内容的提问来源于stack exchange,提问作者Jkind9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 20:42:03