You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Numpy在3D图像数组中定位目标矩形的索引位置?

图像中目标矩形的Numpy定位实现

我需要在一张图像中定位指定的目标矩形,当前图像背景为黑色但后续可能发生变化。之前尝试使用OpenCV的matchTemplate方法,却因背景噪声导致检测结果不准确。现在希望通过Numpy直接判断图像数组中是否包含目标数组,并返回对应的位置索引,目前仅能通过固定颜色匹配获取索引,代码如下:

import cv2
import numpy as np
template = cv2.imread('image.png', cv2.IMREAD_COLOR)
target = cv2.imread('target.png', cv2.IMREAD_COLOR)

index_pos = np.where((template[:,:,0]==116) & (template[:,:,1]==148) & 
(template[:,:,2]==208))

解决方案

核心思路

通过滑动窗口遍历+数组比对的方式,在图像中逐个检查所有能容纳目标矩形的区域,判断该区域是否与目标数组完全一致。

基础实现代码

import cv2
import numpy as np

def locate_target(image, target):
    # 获取图像与目标的尺寸参数
    img_h, img_w, _ = image.shape
    tgt_h, tgt_w, _ = target.shape

    # 遍历所有可能的左上角起始坐标
    for y in range(img_h - tgt_h + 1):
        for x in range(img_w - tgt_w + 1):
            # 截取当前窗口区域
            current_window = image[y:y+tgt_h, x:x+tgt_w]
            # 比对窗口与目标是否完全匹配
            if np.array_equal(current_window, target):
                # 返回左上角和右下角坐标
                return (x, y), (x + tgt_w, y + tgt_h)
    # 未匹配到目标时返回None
    return None

# 读取图像与目标
source_img = cv2.imread('image.png', cv2.IMREAD_COLOR)
target_img = cv2.imread('target.png', cv2.IMREAD_COLOR)

# 执行定位
result = locate_target(source_img, target_img)
if result:
    print(f"目标位置:左上角{result[0]},右下角{result[1]}")
else:
    print("未找到目标矩形")

高效优化版本

针对大尺寸图像,上述双重循环效率较低,可以利用Numpy的as_strided生成滑动窗口视图,减少内存开销并提升比对速度:

import cv2
import numpy as np
from numpy.lib.stride_tricks import as_strided

def locate_target_fast(image, target):
    img_h, img_w, img_c = image.shape
    tgt_h, tgt_w, tgt_c = target.shape

    # 生成所有滑动窗口的视图(无额外内存占用)
    window_strides = image.strides + image.strides[:2]
    window_shape = (img_h - tgt_h + 1, img_w - tgt_w + 1, tgt_h, tgt_w, img_c)
    all_windows = as_strided(image, shape=window_shape, strides=window_strides)

    # 批量比对所有窗口与目标
    match_mask = np.all(all_windows == target, axis=(2, 3, 4))
    # 获取匹配的坐标
    match_y, match_x = np.where(match_mask)

    if len(match_y) > 0:
        x, y = match_x[0], match_y[0]
        return (x, y), (x + tgt_w, y + tgt_h)
    return None

# 使用示例
source_img = cv2.imread('image.png', cv2.IMREAD_COLOR)
target_img = cv2.imread('target.png', cv2.IMREAD_COLOR)
fast_result = locate_target_fast(source_img, target_img)
print(fast_result if fast_result else "未找到目标")

内容的提问来源于stack exchange,提问作者kerz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 12:20:42