TensorFlow目标检测模型:如何获取图像中边界框的实际位置?
如何将TensorFlow目标检测的归一化边界框转换为实际像素位置
没问题,我来帮你搞定这个问题!你打印出来的这些数组是TensorFlow Object Detection API返回的归一化边界框坐标,格式是[y_min, x_min, y_max, x_max],每个值都是0到1之间的比例,对应相对于输入图像高度/宽度的位置。要得到图像中的实际像素位置,只需要把这些比例值乘以图像的实际高度和宽度就可以了。
下面分两种常见场景详细说明:
场景1:输入模型的是原始尺寸图像
如果你的图像没有经过resize,直接送入模型检测,按以下步骤操作:
获取原始图像尺寸
用PIL或OpenCV读取图像后,提取它的高度和宽度。比如用PIL的写法:from PIL import Image img = Image.open("你的图像路径.jpg") width, height = img.size # 注意:PIL返回的是(width, height),而numpy数组的维度是(height, width, channels)转换归一化坐标到实际像素
假设你已经拿到了归一化的boxes数组(就是你打印的结果),直接按对应维度相乘即可:import numpy as np # 示例归一化boxes(你输出的结果片段) normalized_boxes = np.array([ [0.48789895, 0.45768291, 0.97402203, 0.74386948], [0.45094413, 0.43764329, 0.95940024, 0.83383584], # 其他检测框... ]) # 计算实际像素坐标 y_min = normalized_boxes[:, 0] * height # 左上角y坐标 x_min = normalized_boxes[:, 1] * width # 左上角x坐标 y_max = normalized_boxes[:, 2] * height # 右下角y坐标 x_max = normalized_boxes[:, 3] * width # 右下角x坐标 # 组合成常用格式:(x_min, y_min, x_max, y_max) actual_boxes = np.stack([x_min, y_min, x_max, y_max], axis=1) # 如果需要整数像素(比如绘制边界框时),可以转成int类型 actual_boxes = actual_boxes.astype(np.int32)
场景2:输入模型的是resize后的图像
如果为了匹配模型输入要求,你把原始图像resize成了固定尺寸(比如640x640),那需要先把归一化坐标转成resize后图像的像素,再映射回原始图像:
记录原始与resize后的图像尺寸
# 原始图像尺寸 orig_img = Image.open("你的图像路径.jpg") orig_width, orig_height = orig_img.size # resize后的尺寸(比如模型要求的输入尺寸) resized_width, resized_height = 640, 640先转resize后像素,再映射回原始图像
# 转换为resize后的像素坐标 resized_y_min = normalized_boxes[:, 0] * resized_height resized_x_min = normalized_boxes[:, 1] * resized_width resized_y_max = normalized_boxes[:, 2] * resized_height resized_x_max = normalized_boxes[:, 3] * resized_width # 计算缩放比例 scale_x = orig_width / resized_width scale_y = orig_height / resized_height # 映射回原始图像的像素坐标 orig_x_min = resized_x_min * scale_x orig_y_min = resized_y_min * scale_y orig_x_max = resized_x_max * scale_x orig_y_max = resized_y_max * scale_y orig_boxes = np.stack([orig_x_min, orig_y_min, orig_x_max, orig_y_max], axis=1) orig_boxes = orig_boxes.astype(np.int32)
关键注意事项
- 别搞混坐标顺序:TensorFlow返回的boxes是
[y_min, x_min, y_max, x_max],也就是先y轴再x轴,和我们日常习惯的(x, y)顺序相反,这是新手最容易踩坑的点。 - 如果用官方detect_fn:可以直接从检测结果中提取boxes,结合输入图像的numpy数组尺寸计算,示例如下:
import tensorflow as tf # 假设你已经加载了预训练的detect_fn image_np = np.array(Image.open("你的图像路径.jpg")) input_tensor = tf.convert_to_tensor(np.expand_dims(image_np, 0), dtype=tf.float32) detections = detect_fn(input_tensor) # 提取归一化boxes(取第一个batch的结果) normalized_boxes = detections['detection_boxes'][0].numpy() # 获取图像的高度和宽度(numpy数组的shape是(height, width, channels)) height, width, _ = image_np.shape # 转换为实际像素坐标 y_min = normalized_boxes[:, 0] * height x_min = normalized_boxes[:, 1] * width y_max = normalized_boxes[:, 2] * height x_max = normalized_boxes[:, 3] * width
内容的提问来源于stack exchange,提问作者OMAR MEDHAT
相关产品推荐
相关产品推荐

