如何利用归一化边界多边形顶点绘制Bounding Box?
解决Google Vision API归一化顶点绘制Bounding Box问题
问题说明
我正在使用Google Vision API的localization模块,接口返回了normalized bounding polygon vertices,但尝试在目标上绘制bounding boxes时未能成功。以下是接口返回结果和我编写的代码:
API返回结果
Top (confidence: 0.8741532564163208) Normalized bounding polygon vertices: - (0.3563929498195648, 0.36136594414711) - (0.6341143250465393, 0.36136594414711) - (0.6341143250465393, 0.6402543783187866) - (0.3563929498195648, 0.6402543783187866) Luggage & bags (confidence: 0.8460243940353394) Normalized bounding polygon vertices: - (0.5353205800056458, 0.6736522316932678) - (0.6584492921829224, 0.6736522316932678) - (0.6584492921829224, 0.7805569767951965) - (0.5353205800056458, 0.7805569767951965) Shoe (confidence: 0.6873495578765869) Normalized bounding polygon vertices: - (0.001777957659214735, 0.8177978992462158) - (0.10019400715827942, 0.8177978992462158) - (0.10019400715827942, 0.9128244519233704) - (0.001777957659214735, 0.9128244519233704)
我编写的代码
import cv2 import numpy as np # Load the image img = cv2.imread('path to image') # Define the polygon vertices vertices = np.array([(0.08251162618398666, 0.7436794638633728), (0.18944908678531647, 0.7436794638633728), (0.18944908678531647, 0.8542687892913818),(0.08251162618398666, 0.8542687892913818)]) # Convert the normalized vertices to pixel coordinates height, width = img.shape[:2] pixels = np.array([(int(vertex[0] * width), int(vertex[1] * height)) for vertex in vertices]) # Draw the polygon on the image cv2.polylines(img, [pixels], True, (0, 255, 0), 2) # Display the result cv2.imshow('Image with polygon', img) cv2.waitKey(0)
修正方案
下面是整合API返回结果、正确绘制所有目标Bounding Box的代码:
import cv2 import numpy as np # 替换为你的图片路径 img = cv2.imread('your_image_path.jpg') if img is None: print("图像加载失败,请检查路径") exit() height, width = img.shape[:2] # 将API返回的检测结果整理为结构化列表 detections = [ { "label": "Top", "confidence": 0.8741532564163208, "vertices": [(0.3563929498195648, 0.36136594414711), (0.6341143250465393, 0.36136594414711), (0.6341143250465393, 0.6402543783187866), (0.3563929498195648, 0.6402543783187866)] }, { "label": "Luggage & bags", "confidence": 0.8460243940353394, "vertices": [(0.5353205800056458, 0.6736522316932678), (0.6584492921829224, 0.6736522316932678), (0.6584492921829224, 0.7805569767951965), (0.5353205800056458, 0.7805569767951965)] }, { "label": "Shoe", "confidence": 0.6873495578765869, "vertices": [(0.001777957659214735, 0.8177978992462158), (0.10019400715827942, 0.8177978992462158), (0.10019400715827942, 0.9128244519233704), (0.001777957659214735, 0.9128244519233704)] } ] # 遍历所有检测结果,绘制框和标签 for det in detections: # 转换归一化坐标为像素坐标 pixel_vertices = np.array([(int(v[0] * width), int(v[1] * height)) for v in det["vertices"]], np.int32) # 绘制多边形框 cv2.polylines(img, [pixel_vertices], isClosed=True, color=(0, 255, 0), thickness=2) # 添加类别和置信度标签 label = f"{det['label']} ({det['confidence']:.2f})" label_pos = (pixel_vertices[0][0], pixel_vertices[0][1] - 10) cv2.putText(img, label, label_pos, cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 255, 0), 2) # 显示结果 cv2.imshow('Detected Objects', img) cv2.waitKey(0) cv2.destroyAllWindows()
关键修正点
- 直接使用API返回的实际顶点数据,避免手动硬编码错误的坐标
- 增加图像加载失败的判断,提前排查路径问题
- 批量处理所有检测目标,一次性绘制所有Bounding Box
- 添加类别标签和置信度显示,结果更直观
- 确保坐标转换逻辑正确:将0-1范围的归一化坐标,乘以图像实际宽高得到像素坐标
内容的提问来源于stack exchange,提问作者NevinTroy
相关产品推荐
相关产品推荐

