You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用归一化边界多边形顶点绘制Bounding Box?

解决Google Vision API归一化顶点绘制Bounding Box问题

问题说明

我正在使用Google Vision API的localization模块,接口返回了normalized bounding polygon vertices,但尝试在目标上绘制bounding boxes时未能成功。以下是接口返回结果和我编写的代码:

API返回结果

Top (confidence: 0.8741532564163208)
Normalized bounding polygon vertices: 
 - (0.3563929498195648, 0.36136594414711)
 - (0.6341143250465393, 0.36136594414711)
 - (0.6341143250465393, 0.6402543783187866)
 - (0.3563929498195648, 0.6402543783187866)

Luggage & bags (confidence: 0.8460243940353394)
Normalized bounding polygon vertices: 
 - (0.5353205800056458, 0.6736522316932678)
 - (0.6584492921829224, 0.6736522316932678)
 - (0.6584492921829224, 0.7805569767951965)
 - (0.5353205800056458, 0.7805569767951965)

Shoe (confidence: 0.6873495578765869)
Normalized bounding polygon vertices: 
 - (0.001777957659214735, 0.8177978992462158)
 - (0.10019400715827942, 0.8177978992462158)
 - (0.10019400715827942, 0.9128244519233704)
 - (0.001777957659214735, 0.9128244519233704)

我编写的代码

import cv2
import numpy as np

# Load the image
img = cv2.imread('path to image')

# Define the polygon vertices
vertices = np.array([(0.08251162618398666, 0.7436794638633728), (0.18944908678531647, 0.7436794638633728),
                     (0.18944908678531647, 0.8542687892913818),(0.08251162618398666, 0.8542687892913818)])

# Convert the normalized vertices to pixel coordinates
height, width = img.shape[:2]
pixels = np.array([(int(vertex[0] * width), int(vertex[1] * height)) for vertex in vertices])

# Draw the polygon on the image
cv2.polylines(img, [pixels], True, (0, 255, 0), 2)

# Display the result
cv2.imshow('Image with polygon', img)
cv2.waitKey(0)

修正方案

下面是整合API返回结果、正确绘制所有目标Bounding Box的代码:

import cv2
import numpy as np

# 替换为你的图片路径
img = cv2.imread('your_image_path.jpg')
if img is None:
    print("图像加载失败,请检查路径")
    exit()

height, width = img.shape[:2]

# 将API返回的检测结果整理为结构化列表
detections = [
    {
        "label": "Top",
        "confidence": 0.8741532564163208,
        "vertices": [(0.3563929498195648, 0.36136594414711),
                     (0.6341143250465393, 0.36136594414711),
                     (0.6341143250465393, 0.6402543783187866),
                     (0.3563929498195648, 0.6402543783187866)]
    },
    {
        "label": "Luggage & bags",
        "confidence": 0.8460243940353394,
        "vertices": [(0.5353205800056458, 0.6736522316932678),
                     (0.6584492921829224, 0.6736522316932678),
                     (0.6584492921829224, 0.7805569767951965),
                     (0.5353205800056458, 0.7805569767951965)]
    },
    {
        "label": "Shoe",
        "confidence": 0.6873495578765869,
        "vertices": [(0.001777957659214735, 0.8177978992462158),
                     (0.10019400715827942, 0.8177978992462158),
                     (0.10019400715827942, 0.9128244519233704),
                     (0.001777957659214735, 0.9128244519233704)]
    }
]

# 遍历所有检测结果,绘制框和标签
for det in detections:
    # 转换归一化坐标为像素坐标
    pixel_vertices = np.array([(int(v[0] * width), int(v[1] * height)) for v in det["vertices"]], np.int32)
    # 绘制多边形框
    cv2.polylines(img, [pixel_vertices], isClosed=True, color=(0, 255, 0), thickness=2)
    # 添加类别和置信度标签
    label = f"{det['label']} ({det['confidence']:.2f})"
    label_pos = (pixel_vertices[0][0], pixel_vertices[0][1] - 10)
    cv2.putText(img, label, label_pos, cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 255, 0), 2)

# 显示结果
cv2.imshow('Detected Objects', img)
cv2.waitKey(0)
cv2.destroyAllWindows()

关键修正点

  • 直接使用API返回的实际顶点数据,避免手动硬编码错误的坐标
  • 增加图像加载失败的判断,提前排查路径问题
  • 批量处理所有检测目标,一次性绘制所有Bounding Box
  • 添加类别标签和置信度显示,结果更直观
  • 确保坐标转换逻辑正确:将0-1范围的归一化坐标,乘以图像实际宽高得到像素坐标

内容的提问来源于stack exchange,提问作者NevinTroy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 15:37:08