如何将YOLOv8 TensorFlow Lite模型原始输出解码为边界框坐标与类别概率
解码YOLO TFLite输出张量的简便方法
核心步骤
针对你这个[1,7,8400]的输出张量,解码流程可按以下几步完成:
- 去除批次维度,调整张量形状为便于处理的格式
- 分离边界框坐标与类别概率
- 将归一化坐标转换为图像绝对像素坐标
- 用非极大值抑制(NMS)过滤冗余检测框
代码实现
下面是基于NumPy和OpenCV的简洁实现(也可替换为TensorFlow API适配不同部署环境):
import numpy as np import cv2 def decode_yolo_tflite_output(output_tensor, img_size=640, conf_threshold=0.5, iou_threshold=0.45): # 去除批次维度,转置为[8400,7],每行对应一个检测候选框 predictions = np.squeeze(output_tensor).T # shape: (8400,7) # 分离坐标和类别概率 xywh = predictions[:, :4] class_probs = predictions[:, 4:] # 计算每个框的最大类别概率及对应类别索引 confidences = np.max(class_probs, axis=1) class_ids = np.argmax(class_probs, axis=1) # 过滤低置信度的候选框 mask = confidences >= conf_threshold xywh = xywh[mask] confidences = confidences[mask] class_ids = class_ids[mask] # 将归一化的xywh转换为绝对像素坐标的xyxy(左上角、右下角) xyxy = np.copy(xywh) xyxy[:, 0] = xywh[:, 0] * img_size - xywh[:, 2] * img_size / 2 # x1 xyxy[:, 1] = xywh[:, 1] * img_size - xywh[:, 3] * img_size / 2 # y1 xyxy[:, 2] = xywh[:, 0] * img_size + xywh[:, 2] * img_size / 2 # x2 xyxy[:, 3] = xywh[:, 1] * img_size + xywh[:, 3] * img_size / 2 # y3 # 非极大值抑制去除重复框 indices = cv2.dnn.NMSBoxes(xyxy.tolist(), confidences.tolist(), conf_threshold, iou_threshold) # 提取最终有效检测结果 final_boxes = xyxy[indices] final_confidences = confidences[indices] final_class_ids = class_ids[indices] return final_boxes, final_confidences, final_class_ids # 使用示例 # 假设output是你的TFLite模型输出张量,形状为[1,7,8400] boxes, confs, class_ids = decode_yolo_tflite_output(output)
说明
- 若需适配TensorFlow部署环境,可将NumPy操作替换为
tf.squeeze、tf.transpose、tf.image.non_max_suppression等API conf_threshold(置信度阈值)和iou_threshold(重叠度阈值)可根据你的数据集调整,优化检测精度- 坐标转换基于YOLO输出为相对输入图像尺寸的归一化值的前提,与你640x640的输入图像完全匹配
内容的提问来源于stack exchange,提问作者Santa Bot
相关产品推荐
相关产品推荐

