TFLite版YOLOv5推理报错无法解析输出获取检测框
问题背景
基于YOLOv5训练了可检测图像中+字符的检测模型,转换为TFLite格式部署时,执行图像推理无法正确解析模型输出,触发IndexError: list index out of range报错。
现有推理实现代码
interpreter = tf.lite.Interpreter("/Users/maximereder/Desktop/best-fp16.tflite") interpreter.allocate_tensors() IMAGE_PATH = "/Users/maximereder/Documents/ML/dataset/plus-1000x750/train/IMG_1492B2.jpg" img = cv2.resize(cv2.imread(IMAGE_PATH), (640, 640)) features = img.copy() np_features = np.array(features, dtype=np.float32) np_features = np.expand_dims(np_features, axis=0) input_details = interpreter.get_input_details() output_details = interpreter.get_output_details() interpreter.set_tensor(input_details[0]['index'], np_features) interpreter.invoke()
输出节点信息
查询得到模型仅包含1个输出节点,信息如下:
[{'name': 'Identity', 'index': 422, 'shape': array([ 1, 25200, 85], dtype=int32), 'shape_signature': array([ 1, 25200, 85], dtype=int32), 'dtype': numpy.float32, 'quantization': (0.0, 0), 'quantization_parameters': {'scales': array([], dtype=float32), 'zero_points': array([], dtype=int32), 'quantized_dimension': 0}, 'sparsity_parameters': {}}]
读取output_details[0]可获得形状为[1,25200,85]的张量,样例数据如下:
array([[[-5.6185187e-03, 1.3539949e-02, 5.7405889e-02, 4.2354122e-02, 1.3114554e-04, 9.9999905e-01], [-4.4679684e-03, 1.7201375e-02, 6.8576269e-02, 2.5891241e-02, 3.4220223e-04, 9.9999964e-01], [-4.9383980e-03, 1.5453462e-02, 4.2031817e-02, 2.6558569e-02, 2.5974249e-03, 9.9999815e-01], ..., [ 9.2678040e-01, 9.2856336e-01, 7.1995622e-01, 5.6132025e-01, 1.4253161e-14, 9.9999440e-01], [ 9.2535079e-01, 9.2647862e-01, 9.4501650e-01, 1.1292257e+00, 4.9554409e-09, 9.9999994e-01], [ 1.0224271e+00, 9.7982901e-01, 2.2890522e+00, 1.1467136e-02, 1.4553191e-07, 9.9999893e-01]]], dtype=float32)
报错触发逻辑
按照常规多输出目标检测模型的解析逻辑,尝试分别读取检测框、检测类别、置信度、检测框数量:
detection_boxes = interpreter.get_tensor(output_details[0]['index']) detection_classes = interpreter.get_tensor(output_details[1]['index']) detection_scores = interpreter.get_tensor(output_details[2]['index']) num_boxes = interpreter.get_tensor(output_details[3]['index'])
执行时触发索引越界报错,不清楚单输出张量的结构,无法正确提取检测结果。
问题原因与解决方法
报错根因
你导出的YOLOv5 TFLite是合并单输出格式,和TensorFlow官方预训练检测模型的4输出结构完全不同,直接按4节点索引读取必然越界。
形状为[1,25200,85]是YOLOv5 640x640输入下的标准原始检测输出:
- 第一维
1对应batch size - 第二维
25200对应三个检测尺度的总锚框数(80803 + 40403 + 20203 = 25200) - 第三维
85对应每个锚框的预测值,顺序为[x_center, y_center, width, height, 目标置信度, 类别1置信度, 类别2置信度...类别80置信度]。你是单类别检测(仅检测+),所以只有第6位(索引5)的类别置信度是有效值,后续79位都是无效占位值。
另外你的现有预处理存在错误:YOLOv5默认要求输入像素值归一化到01区间,直接传入0255范围的BGR像素会导致预测结果完全失效。
正确解析步骤
- 修正预处理逻辑,对齐训练时的预处理规则(通常为RGB通道、0~1像素归一化)
- 读取唯一的输出张量,拆分每个锚框的坐标、目标置信度、类别置信度
- 计算最终检测置信度(目标置信度 * 类别置信度),过滤低于置信度阈值的锚框
- 将xywh格式的中心坐标转换为xyxy格式的角点坐标,映射回原图尺寸
- 执行非极大值抑制(NMS)去除重复检测框,得到最终结果
参考实现代码:
import cv2 import numpy as np import tensorflow as tf # 超参数配置 CONF_THRESHOLD = 0.25 IOU_THRESHOLD = 0.45 INPUT_SIZE = 640 # 初始化模型 interpreter = tf.lite.Interpreter("/Users/maximereder/Desktop/best-fp16.tflite") interpreter.allocate_tensors() input_info = interpreter.get_input_details() output_info = interpreter.get_output_details() # 数据预处理 img = cv2.imread("/Users/maximereder/Documents/ML/dataset/plus-1000x750/train/IMG_1492B2.jpg") orig_h, orig_w = img.shape[:2] img_resized = cv2.resize(img, (INPUT_SIZE, INPUT_SIZE)) img_rgb = cv2.cvtColor(img_resized, cv2.COLOR_BGR2RGB) input_tensor = np.expand_dims(np.array(img_rgb, dtype=np.float32) / 255.0, axis=0) # 执行推理 interpreter.set_tensor(input_info[0]['index'], input_tensor) interpreter.invoke() preds = interpreter.get_tensor(output_info[0]['index'])[0] # 去除batch维度,形状变为[25200, 85] # 解析检测结果 candidate_boxes = [] candidate_scores = [] for pred in preds: x_c, y_c, w, h, obj_conf = pred[:5] cls_conf = pred[5] # 单类别取第一个类别置信度 score = obj_conf * cls_conf if score < CONF_THRESHOLD: continue # 坐标转换映射回原图 x1 = int((x_c - w/2) / INPUT_SIZE * orig_w) y1 = int((y_c - h/2) / INPUT_SIZE * orig_h) x2 = int((x_c + w/2) / INPUT_SIZE * orig_w) y2 = int((y_c + h/2) / INPUT_SIZE * orig_h) # 裁剪越界坐标 x1, y1 = max(0, x1), max(0, y1) x2, y2 = min(orig_w, x2), min(orig_h, y2) candidate_boxes.append([x1, y1, x2, y2]) candidate_scores.append(float(score)) # NMS去重 final_results = [] if candidate_boxes: candidate_boxes = np.array(candidate_boxes) candidate_scores = np.array(candidate_scores) keep_idx = cv2.dnn.NMSBoxes( candidate_boxes.tolist(), candidate_scores.tolist(), CONF_THRESHOLD, IOU_THRESHOLD ).flatten() final_results = [ {"box": candidate_boxes[i], "score": candidate_scores[i]} for i in keep_idx ] print(f"检测到{len(final_results)}个'+'字符") else: print("未检测到目标")
注意:如果运行后发现坐标偏移,可检查坐标是否为640尺寸下的绝对像素值(坐标范围0~640),如果是则不需要除以
INPUT_SIZE,直接按比例映射回原图即可。官方导出脚本默认会把坐标解码逻辑合并到模型中,无需额外实现Sigmoid、锚框换算步骤。
内容的提问来源于stack exchange,提问作者Maxime R.
相关产品推荐
相关产品推荐

