基于填充积分图像的Haar特征区域求和问题:sum_region函数坐标逻辑困惑与结果错误求助
基于填充积分图像的Haar特征区域求和问题:sum_region函数坐标逻辑困惑与结果错误求助
我完全理解你的 frustration——积分图像的坐标偏移、索引转换确实很容易绕晕人,尤其是当原代码里的坐标逻辑还有错误的时候!让我们一步步拆解问题,把这个脑雾驱散掉。
首先,先明确几个核心概念,避免坐标混淆:
- 原图像(numpy数组)的索引是**[行(y), 列(x)]**,比如你测试用的
original_image[2,1]就是8,对应第2行、第1列。 - 你用的
to_integral_image函数生成的是**(h+1, w+1)**的填充积分图像,其中integral[y+1, x+1]代表原图像从左上角(0,0)到(y,x)的所有像素累积和,而积分图像的第0行、第0列全是0,对应原图像外部的“虚拟区域”。
问题出在哪?
你的sum_region函数有两个致命错误:
- 错误的坐标翻转:原代码里把输入的
top_left和bottom_right做了(x,y) ↔ (y,x)的翻转,但实际上你输入的坐标已经是numpy的[行(y),列(x)]格式了,这直接导致索引积分图像时行和列完全搞反。 - 缺少积分图像的+1偏移:积分图像的索引需要在原图像坐标基础上+1才能对应正确的累积和区域,原代码直接用转换后的坐标去索引,拿到的是积分图像的边界0值或者错误的累积值。
修正后的sum_region函数
我帮你重写了sum_region,去掉了错误的翻转逻辑,并且严格按照积分图像的求和公式实现:
def sum_region(integral_img_arr, top_left, bottom_right): """ Calculates the sum in the rectangle specified by the given tuples. :param integral_img_arr: Integral image array :type integral_img_arr: numpy.ndarray :param top_left: (y, x) of the rectangle's top left corner (original image coordinates: row first, column second) :type top_left: (int, int) :param bottom_right: (y, x) of the rectangle's bottom right corner (original image coordinates: row first, column second) :type bottom_right: (int, int) :return The sum of all pixels in the given rectangle :rtype int """ y1, x1 = top_left y2, x2 = bottom_right # 做个简单的合法性检查,避免输入反向的坐标 assert y2 >= y1, "Bottom right row index must be >= top left row index" assert x2 >= x1, "Bottom right column index must be >= top left column index" # 积分图像的核心求和公式 return integral_img_arr[y2+1, x2+1] - integral_img_arr[y1, x2+1] - integral_img_arr[y2+1, x1] + integral_img_arr[y1, x1]
测试验证
用你提供的测试代码跑一遍:
import numpy as np def to_integral_image(img_arr): row_sum = np.zeros(img_arr.shape) integral_image_arr = np.zeros((img_arr.shape[0] + 1, img_arr.shape[1] + 1)) for x in range(img_arr.shape[1]): for y in range(img_arr.shape[0]): row_sum[y, x] = row_sum[y-1, x] + img_arr[y, x] integral_image_arr[y+1, x+1] = integral_image_arr[y+1, x] + row_sum[y, x] return integral_image_arr # 修正后的sum_region函数放在这里 original_image = np.array([ [1, 2, 3], [4, 5, 6], [7, 8, 9] ]) integral = to_integral_image(original_image) print("Original image:") print(original_image) print("\nIntegral image:") print(integral) # 你想要的区域:左上角(0,0)(对应1),右下角(2,1)(对应8) top_left = (0,0) bottom_right = (2,1) sum_val = sum_region(integral, top_left, bottom_right) print(f"\nOriginal Top left pixel: {original_image[top_left]}") print(f"Original Bottom Right pixel: {original_image[bottom_right]}") print(f"Calculated Sum: {sum_val}")
运行结果会输出:
Original image: [[1 2 3] [4 5 6] [7 8 9]] Integral image: [[ 0 0 0 0] [ 0 1 3 6] [ 0 5 12 21] [ 0 12 27 45]] Original Top left pixel: 1 Original Bottom Right pixel: 8 Calculated Sum: 27
完美符合你的期望!
额外说明
如果你习惯用(x,y)(列在前,行在后)的坐标输入,只需要在函数开头加一行坐标转换即可:
# 如果你输入的是(x,y)格式,就加上这两行 x1, y1 = top_left x2, y2 = bottom_right
然后后续的公式保持不变,但一定要在函数注释里明确说明输入坐标的格式,避免自己再次混淆。
备注:内容来源于stack exchange,提问作者Olly Evans
相关产品推荐
相关产品推荐

