二进制图像围绕多数白色像素生成正确边界框的代码问题
问题描述
给定一张二进制图像(例如Canny边缘分割输出的图像),希望围绕多数白色像素绘制边界框。使用np.nonzero定位白色像素点,并编写了get_bounding_box函数获取边界框坐标,但返回的左上角、右下角坐标不符合预期。
相关代码如下:
def get_bounding_box(image,thresh=0.95): nonzero_indices = np.nonzero(image) min_row, max_row = np.min(nonzero_indices[0]), np.max(nonzero_indices[0]) min_col, max_col = np.min(nonzero_indices[1]), np.max(nonzero_indices[1]) box_size = max_row - min_row + 1, max_col - min_col + 1 print(box_size) #box_size_thresh = (int(box_size[0] * thresh), int(box_size[1] * thresh)) box_size_thresh = (int(box_size[0]), int(box_size[1])) #coordinates of the box that contains 95% of the highest pixel values top_left = (min_row + int((box_size[0] - box_size_thresh[0]) / 2), min_col + int((box_size[1] - box_size_thresh[1]) / 2)) bottom_right = (top_left[0] + box_size_thresh[0], top_left[1] + box_size_thresh[1]) print((top_left[0], top_left[1]), (bottom_right[0], bottom_right[1])) return (top_left[0], top_left[1]), (bottom_right[0], bottom_right[1])
调用代码:
seg= canny_segmentation(gray) bb_thresh = get_bounding_box(seg,0.95) im_crop = gray[bb_thresh[0][1]:bb_thresh[1][1],bb_thresh[0][0]:bb_thresh[1][0]]
问题原因与解决办法
1. 坐标顺序完全颠倒
NumPy数组的索引逻辑是[行, 列],对应图像的[y轴坐标, x轴坐标],但你在裁剪图像时完全搞反了坐标的使用顺序:
- 函数返回的
top_left是(min_row, min_col),对应图像的(y1, x1),bottom_right是(y2, x2) - 而裁剪代码写成了
gray[bb_thresh[0][1]:bb_thresh[1][1],bb_thresh[0][0]:bb_thresh[1][0]],把x轴坐标放在了行索引的位置,y轴坐标放在了列索引的位置,直接导致裁剪区域错位。
2. 阈值逻辑未生效
你注释掉了实现“保留95%像素区域”的核心代码box_size_thresh = (int(box_size[0] * thresh), int(box_size[1] * thresh)),改用了原边界框尺寸,这会让函数直接返回最外层的全量边界框,无法实现“围绕多数像素”的目标。
修正后的代码
修正边界框函数
def get_bounding_box(image, thresh=0.95): nonzero_indices = np.nonzero(image) min_row, max_row = np.min(nonzero_indices[0]), np.max(nonzero_indices[0]) min_col, max_col = np.min(nonzero_indices[1]), np.max(nonzero_indices[1]) box_size = max_row - min_row + 1, max_col - min_col + 1 # 恢复阈值计算,生成包含95%像素的目标框尺寸 box_size_thresh = (int(box_size[0] * thresh), int(box_size[1] * thresh)) # 计算中心偏移量,确保目标框在原边界框内居中 offset_row = (box_size[0] - box_size_thresh[0]) // 2 offset_col = (box_size[1] - box_size_thresh[1]) // 2 top_left = (min_row + offset_row, min_col + offset_col) bottom_right = (top_left[0] + box_size_thresh[0], top_left[1] + box_size_thresh[1]) return top_left, bottom_right
修正调用与裁剪代码
seg = canny_segmentation(gray) bb_thresh = get_bounding_box(seg, 0.95) # 按照[y1:y2, x1:x2]的正确索引顺序裁剪 im_crop = gray[bb_thresh[0][0]:bb_thresh[1][0], bb_thresh[0][1]:bb_thresh[1][1]]
额外优化建议
如果Canny输出是浮点型图像(值在0-1之间),np.nonzero会把所有大于0的像素纳入计算,可能包含噪点。可以先做二值化处理:
# 只保留值大于0.5的有效边缘像素,过滤噪点 seg = (seg > 0.5).astype(np.uint8)
内容的提问来源于stack exchange,提问作者zaza
相关产品推荐
相关产品推荐

