如何在Python中将Azure Custom Vision标注格式转换为YOLO v8格式
Azure Custom Vision 边界框格式转YOLO/OpenCV兼容格式解决方案
问题核心
通过Azure Custom Vision API获取车牌检测结果时,返回的边界框采用相对坐标格式{left, top, width, height},而OpenCV裁剪、YOLO v8使用的是绝对坐标格式(x1, y1, x2, y2)(x1/y1为左上角坐标,x2/y2为右下角坐标)。原代码因裁剪时坐标顺序错误,导致提取的车牌区域异常。
格式转换公式
先将Azure返回的相对坐标转换为图像的绝对坐标:
x1 = left * image_width(左上角x坐标)y1 = top * image_height(左上角y坐标)x2 = x1 + width * image_width(右下角x坐标)y2 = y1 + height * image_height(右下角y坐标)
OpenCV图像裁剪的正确语法为:图像[y1:y2, x1:x2],需注意先指定y轴范围,后指定x轴范围。
修正后的完整代码
import requests import json import cv2 # 调用Azure Custom Vision API获取检测结果 url = "你的Custom Vision模型预测URL" headers = {'content-type': 'application/octet-stream'} # 适配二进制图像上传的请求头 with open("你的图像路径.jpg", "rb") as image_file: r = requests.post(url, data=image_file, headers=headers) # 解析API返回的JSON结果 pred = json.loads(r.content) image_path = "你的图像路径.jpg" image = cv2.imread(image_path) image_height, image_width, _ = image.shape # 遍历检测结果,转换格式并处理 for prediction in pred["predictions"]: # 过滤低置信度的检测结果 if prediction["probability"] > 0.5: bbox = prediction['boundingBox'] # 转换为图像绝对坐标 x1 = int(bbox['left'] * image_width) y1 = int(bbox['top'] * image_height) x2 = int(x1 + bbox['width'] * image_width) y2 = int(y1 + bbox['height'] * image_height) # 在原图上绘制边界框 cv2.rectangle(image, (x1, y1), (x2, y2), (0, 0, 255), 5) # 正确裁剪车牌区域 license_plate_crop = image[y1:y2, x1:x2, :] # 显示裁剪后的车牌 cv2.imshow('裁剪的车牌', license_plate_crop) cv2.waitKey(0) # 显示带边界框的原图 cv2.imshow('检测结果', image) cv2.waitKey(0) cv2.destroyAllWindows()
关键修正点
- 修复裁剪坐标顺序:原代码错误使用
image[int(left):int(top), ...],改为符合OpenCV规则的image[y1:y2, x1:x2] - 统一变量存储转换后的坐标,避免重复计算
- 增加窗口资源释放逻辑,防止程序结束后窗口残留
额外说明
- Azure Custom Vision返回的
left/top/width/height均为相对于图像宽高的比例值(范围0~1),必须乘以图像实际宽高才能得到绝对坐标 - 转换后的
x1,y1,x2,y2格式可直接用于YOLO v8的训练数据标注
内容的提问来源于stack exchange,提问作者Shahad
相关产品推荐
相关产品推荐

