You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

二进制图像围绕多数白色像素生成正确边界框的代码问题

问题描述

给定一张二进制图像(例如Canny边缘分割输出的图像),希望围绕多数白色像素绘制边界框。使用np.nonzero定位白色像素点,并编写了get_bounding_box函数获取边界框坐标,但返回的左上角、右下角坐标不符合预期。

相关代码如下:

def get_bounding_box(image,thresh=0.95):
    nonzero_indices = np.nonzero(image)
    min_row, max_row = np.min(nonzero_indices[0]), np.max(nonzero_indices[0])
    min_col, max_col = np.min(nonzero_indices[1]), np.max(nonzero_indices[1])
    box_size = max_row - min_row + 1, max_col - min_col + 1
    print(box_size)
    #box_size_thresh = (int(box_size[0] * thresh), int(box_size[1] * thresh))
    box_size_thresh = (int(box_size[0]), int(box_size[1]))
    #coordinates of the box that contains 95% of the highest pixel values
    top_left = (min_row + int((box_size[0] - box_size_thresh[0]) / 2), min_col + int((box_size[1] - box_size_thresh[1]) / 2))
    bottom_right = (top_left[0] + box_size_thresh[0], top_left[1] + box_size_thresh[1])
    print((top_left[0], top_left[1]), (bottom_right[0], bottom_right[1]))
    return (top_left[0], top_left[1]), (bottom_right[0], bottom_right[1])

调用代码:

seg= canny_segmentation(gray)
bb_thresh = get_bounding_box(seg,0.95)
im_crop = gray[bb_thresh[0][1]:bb_thresh[1][1],bb_thresh[0][0]:bb_thresh[1][0]]  
问题原因与解决办法

1. 坐标顺序完全颠倒

NumPy数组的索引逻辑是[行, 列],对应图像的[y轴坐标, x轴坐标],但你在裁剪图像时完全搞反了坐标的使用顺序:

  • 函数返回的top_left是(min_row, min_col),对应图像的(y1, x1),bottom_right是(y2, x2)
  • 而裁剪代码写成了gray[bb_thresh[0][1]:bb_thresh[1][1],bb_thresh[0][0]:bb_thresh[1][0]],把x轴坐标放在了行索引的位置,y轴坐标放在了列索引的位置,直接导致裁剪区域错位。

2. 阈值逻辑未生效

你注释掉了实现“保留95%像素区域”的核心代码box_size_thresh = (int(box_size[0] * thresh), int(box_size[1] * thresh)),改用了原边界框尺寸,这会让函数直接返回最外层的全量边界框,无法实现“围绕多数像素”的目标。

修正后的代码

修正边界框函数

def get_bounding_box(image, thresh=0.95):
    nonzero_indices = np.nonzero(image)
    min_row, max_row = np.min(nonzero_indices[0]), np.max(nonzero_indices[0])
    min_col, max_col = np.min(nonzero_indices[1]), np.max(nonzero_indices[1])
    box_size = max_row - min_row + 1, max_col - min_col + 1
    
    # 恢复阈值计算,生成包含95%像素的目标框尺寸
    box_size_thresh = (int(box_size[0] * thresh), int(box_size[1] * thresh))
    
    # 计算中心偏移量,确保目标框在原边界框内居中
    offset_row = (box_size[0] - box_size_thresh[0]) // 2
    offset_col = (box_size[1] - box_size_thresh[1]) // 2
    top_left = (min_row + offset_row, min_col + offset_col)
    bottom_right = (top_left[0] + box_size_thresh[0], top_left[1] + box_size_thresh[1])
    
    return top_left, bottom_right

修正调用与裁剪代码

seg = canny_segmentation(gray)
bb_thresh = get_bounding_box(seg, 0.95)
# 按照[y1:y2, x1:x2]的正确索引顺序裁剪
im_crop = gray[bb_thresh[0][0]:bb_thresh[1][0], bb_thresh[0][1]:bb_thresh[1][1]]

额外优化建议

如果Canny输出是浮点型图像(值在0-1之间),np.nonzero会把所有大于0的像素纳入计算,可能包含噪点。可以先做二值化处理:

# 只保留值大于0.5的有效边缘像素,过滤噪点
seg = (seg > 0.5).astype(np.uint8)

内容的提问来源于stack exchange,提问作者zaza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 12:51:17