如何将OpenVINO模型服务器的YOLO输出转为检测框并可视化目标?
处理YOLO模型输出并绘制检测框的步骤
你的模型输出形状(1, 84, 8400)对应YOLOv8的标准格式:1个输入样本、84个参数(4个框坐标+80个COCO类别概率)、8400个预测框。以下是完整的处理和绘制流程:
1. 解析输出张量
先去除冗余维度并调整张量形状,方便逐框处理:
import numpy as np # 去掉batch维度,转置为(8400, 84),每行对应一个预测框的参数 output = np.squeeze(output) output = output.T
2. 分离坐标与类别概率
从每个框的参数中拆分出归一化坐标和类别概率:
# 提取归一化的中心坐标(x_center, y_center)、宽(w)、高(h) boxes = output[:, :4] # 提取80个类别的预测概率 class_probs = output[:, 4:]
3. 坐标转换(归一化→像素坐标)
YOLO输出的坐标是相对于原图尺寸的归一化值,需要转换为图片实际像素的左上角/右下角坐标:
import cv2 # 读取原图获取尺寸 img = cv2.imread("zebra.jpeg") h, w = img.shape[:2] # 转换坐标格式 x_center, y_center, box_w, box_h = boxes[:, 0], boxes[:, 1], boxes[:, 2], boxes[:, 3] xmin = (x_center - box_w/2) * w ymin = (y_center - box_h/2) * h xmax = (x_center + box_w/2) * w ymax = (y_center + box_h/2) * h # 组合成(xmin, ymin, xmax, ymax)格式 detections = np.stack([xmin, ymin, xmax, ymax], axis=1)
4. 非极大值抑制(NMS)过滤冗余框
过滤掉重叠度高的低置信度框,保留最可靠的检测结果:
# 获取每个框的最高置信度和对应类别ID confidences = np.max(class_probs, axis=1) class_ids = np.argmax(class_probs, axis=1) # 执行NMS:置信度阈值0.5,IOU重叠阈值0.5 indices = cv2.dnn.NMSBoxes(detections.tolist(), confidences.tolist(), 0.5, 0.5)
5. 加载COCO类别名称
注意你之前导入的imagenet_classes是ImageNet类别,YOLO用的是COCO 80类,替换为正确的类别列表:
coco_classes = [ "person", "bicycle", "car", "motorcycle", "airplane", "bus", "train", "truck", "boat", "traffic light", "fire hydrant", "stop sign", "parking meter", "bench", "bird", "cat", "dog", "horse", "sheep", "cow", "elephant", "bear", "zebra", "giraffe", "backpack", "umbrella", "handbag", "tie", "suitcase", "frisbee", "skis", "snowboard", "sports ball", "kite", "baseball bat", "baseball glove", "skateboard", "surfboard", "tennis racket", "bottle", "wine glass", "cup", "fork", "knife", "spoon", "bowl", "banana", "apple", "sandwich", "orange", "broccoli", "carrot", "hot dog", "pizza", "donut", "cake", "chair", "couch", "potted plant", "bed", "dining table", "toilet", "tv", "laptop", "mouse", "remote", "keyboard", "cell phone", "microwave", "oven", "toaster", "sink", "refrigerator", "book", "clock", "vase", "scissors", "teddy bear", "hair drier", "toothbrush" ]
6. 绘制检测框与类别信息
遍历筛选后的框,在原图上绘制矩形框并标注类别和置信度:
# 遍历NMS筛选后的结果 for i in indices: i = i[0] if isinstance(i, (list, np.ndarray)) else i x1, y1, x2, y2 = detections[i] conf = confidences[i] class_name = coco_classes[class_ids[i]] # 绘制红色矩形框 cv2.rectangle(img, (int(x1), int(y1)), (int(x2), int(y2)), (0, 0, 255), 2) # 绘制类别和置信度文本 text = f"{class_name}: {conf:.2f}" cv2.putText(img, text, (int(x1), int(y1)-10), cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 0, 255), 2) # 保存结果并显示 cv2.imwrite("detection_result.jpg", img) cv2.imshow("YOLO Detection", img) cv2.waitKey(0) cv2.destroyAllWindows()
完整整合代码
把以上步骤整合到你的原有代码中:
import numpy as np import cv2 from ovmsclient import make_grpc_client # COCO 80类名称列表 coco_classes = [ "person", "bicycle", "car", "motorcycle", "airplane", "bus", "train", "truck", "boat", "traffic light", "fire hydrant", "stop sign", "parking meter", "bench", "bird", "cat", "dog", "horse", "sheep", "cow", "elephant", "bear", "zebra", "giraffe", "backpack", "umbrella", "handbag", "tie", "suitcase", "frisbee", "skis", "snowboard", "sports ball", "kite", "baseball bat", "baseball glove", "skateboard", "surfboard", "tennis racket", "bottle", "wine glass", "cup", "fork", "knife", "spoon", "bowl", "banana", "apple", "sandwich", "orange", "broccoli", "carrot", "hot dog", "pizza", "donut", "cake", "chair", "couch", "potted plant", "bed", "dining table", "toilet", "tv", "laptop", "mouse", "remote", "keyboard", "cell phone", "microwave", "oven", "toaster", "sink", "refrigerator", "book", "clock", "vase", "scissors", "teddy bear", "hair drier", "toothbrush" ] # 连接OVMS客户端 client = make_grpc_client("localhost:9000") # 读取图片并预测 with open("zebra.jpeg", "rb") as f: img_bytes = f.read() output = client.predict({"images": img_bytes}, "yolo") # 解析输出 output = np.squeeze(output) output = output.T # 分离坐标和类别概率 boxes = output[:, :4] class_probs = output[:, 4:] # 获取原图尺寸 img = cv2.imread("zebra.jpeg") h, w = img.shape[:2] # 坐标转换 x_center, y_center, box_w, box_h = boxes[:, 0], boxes[:, 1], boxes[:, 2], boxes[:, 3] xmin = (x_center - box_w/2) * w ymin = (y_center - box_h/2) * h xmax = (x_center + box_w/2) * w ymax = (y_center + box_h/2) * h detections = np.stack([xmin, ymin, xmax, ymax], axis=1) # 计算置信度和类别ID confidences = np.max(class_probs, axis=1) class_ids = np.argmax(class_probs, axis=1) # NMS过滤 indices = cv2.dnn.NMSBoxes(detections.tolist(), confidences.tolist(), 0.5, 0.5) # 绘制检测结果 for i in indices: i = i[0] if isinstance(i, (list, np.ndarray)) else i x1, y1, x2, y2 = detections[i] conf = confidences[i] class_name = coco_classes[class_ids[i]] cv2.rectangle(img, (int(x1), int(y1)), (int(x2), int(y2)), (0, 0, 255), 2) text = f"{class_name}: {conf:.2f}" cv2.putText(img, text, (int(x1), int(y1)-10), cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 0, 255), 2) # 保存并显示结果 cv2.imwrite("detection_result.jpg", img) cv2.imshow("Detection Result", img) cv2.waitKey(0) cv2.destroyAllWindows()
内容的提问来源于stack exchange,提问作者Canputer
相关产品推荐
相关产品推荐

