You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TFLite版YOLOv5推理报错无法解析输出获取检测框

问题背景

基于YOLOv5训练了可检测图像中+字符的检测模型,转换为TFLite格式部署时,执行图像推理无法正确解析模型输出,触发IndexError: list index out of range报错。

现有推理实现代码

interpreter = tf.lite.Interpreter("/Users/maximereder/Desktop/best-fp16.tflite")
interpreter.allocate_tensors()

IMAGE_PATH = "/Users/maximereder/Documents/ML/dataset/plus-1000x750/train/IMG_1492B2.jpg"
img = cv2.resize(cv2.imread(IMAGE_PATH), (640, 640))
features = img.copy()
np_features = np.array(features, dtype=np.float32)
np_features = np.expand_dims(np_features, axis=0)

input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()
interpreter.set_tensor(input_details[0]['index'], np_features)

interpreter.invoke()

输出节点信息

查询得到模型仅包含1个输出节点,信息如下:

[{'name': 'Identity',
'index': 422,
'shape': array([    1, 25200,    85], dtype=int32),
'shape_signature': array([    1, 25200,    85], dtype=int32),
'dtype': numpy.float32,
'quantization': (0.0, 0),
'quantization_parameters': {'scales': array([], dtype=float32),
'zero_points': array([], dtype=int32),
'quantized_dimension': 0},
'sparsity_parameters': {}}]

读取output_details[0]可获得形状为[1,25200,85]的张量,样例数据如下:

array([[[-5.6185187e-03,  1.3539949e-02,  5.7405889e-02,  4.2354122e-02,
          1.3114554e-04,  9.9999905e-01],
        [-4.4679684e-03,  1.7201375e-02,  6.8576269e-02,  2.5891241e-02,
          3.4220223e-04,  9.9999964e-01],
        [-4.9383980e-03,  1.5453462e-02,  4.2031817e-02,  2.6558569e-02,
          2.5974249e-03,  9.9999815e-01],
        ...,
        [ 9.2678040e-01,  9.2856336e-01,  7.1995622e-01,  5.6132025e-01,
          1.4253161e-14,  9.9999440e-01],
        [ 9.2535079e-01,  9.2647862e-01,  9.4501650e-01,  1.1292257e+00,
          4.9554409e-09,  9.9999994e-01],
        [ 1.0224271e+00,  9.7982901e-01,  2.2890522e+00,  1.1467136e-02,
          1.4553191e-07,  9.9999893e-01]]], dtype=float32)

报错触发逻辑

按照常规多输出目标检测模型的解析逻辑,尝试分别读取检测框、检测类别、置信度、检测框数量:

detection_boxes = interpreter.get_tensor(output_details[0]['index'])
detection_classes = interpreter.get_tensor(output_details[1]['index'])
detection_scores = interpreter.get_tensor(output_details[2]['index'])
num_boxes = interpreter.get_tensor(output_details[3]['index'])

执行时触发索引越界报错,不清楚单输出张量的结构,无法正确提取检测结果。


问题原因与解决方法

报错根因

你导出的YOLOv5 TFLite是合并单输出格式,和TensorFlow官方预训练检测模型的4输出结构完全不同,直接按4节点索引读取必然越界。
形状为[1,25200,85]是YOLOv5 640x640输入下的标准原始检测输出:

  • 第一维1对应batch size
  • 第二维25200对应三个检测尺度的总锚框数(80803 + 40403 + 20203 = 25200)
  • 第三维85对应每个锚框的预测值,顺序为[x_center, y_center, width, height, 目标置信度, 类别1置信度, 类别2置信度...类别80置信度]。你是单类别检测(仅检测+),所以只有第6位(索引5)的类别置信度是有效值,后续79位都是无效占位值。
    另外你的现有预处理存在错误:YOLOv5默认要求输入像素值归一化到01区间,直接传入0255范围的BGR像素会导致预测结果完全失效。

正确解析步骤

  1. 修正预处理逻辑,对齐训练时的预处理规则(通常为RGB通道、0~1像素归一化)
  2. 读取唯一的输出张量,拆分每个锚框的坐标、目标置信度、类别置信度
  3. 计算最终检测置信度(目标置信度 * 类别置信度),过滤低于置信度阈值的锚框
  4. 将xywh格式的中心坐标转换为xyxy格式的角点坐标,映射回原图尺寸
  5. 执行非极大值抑制(NMS)去除重复检测框,得到最终结果

参考实现代码:

import cv2
import numpy as np
import tensorflow as tf

# 超参数配置
CONF_THRESHOLD = 0.25
IOU_THRESHOLD = 0.45
INPUT_SIZE = 640

# 初始化模型
interpreter = tf.lite.Interpreter("/Users/maximereder/Desktop/best-fp16.tflite")
interpreter.allocate_tensors()
input_info = interpreter.get_input_details()
output_info = interpreter.get_output_details()

# 数据预处理
img = cv2.imread("/Users/maximereder/Documents/ML/dataset/plus-1000x750/train/IMG_1492B2.jpg")
orig_h, orig_w = img.shape[:2]
img_resized = cv2.resize(img, (INPUT_SIZE, INPUT_SIZE))
img_rgb = cv2.cvtColor(img_resized, cv2.COLOR_BGR2RGB)
input_tensor = np.expand_dims(np.array(img_rgb, dtype=np.float32) / 255.0, axis=0)

# 执行推理
interpreter.set_tensor(input_info[0]['index'], input_tensor)
interpreter.invoke()
preds = interpreter.get_tensor(output_info[0]['index'])[0]  # 去除batch维度,形状变为[25200, 85]

# 解析检测结果
candidate_boxes = []
candidate_scores = []
for pred in preds:
    x_c, y_c, w, h, obj_conf = pred[:5]
    cls_conf = pred[5]  # 单类别取第一个类别置信度
    score = obj_conf * cls_conf
    if score < CONF_THRESHOLD:
        continue
    # 坐标转换映射回原图
    x1 = int((x_c - w/2) / INPUT_SIZE * orig_w)
    y1 = int((y_c - h/2) / INPUT_SIZE * orig_h)
    x2 = int((x_c + w/2) / INPUT_SIZE * orig_w)
    y2 = int((y_c + h/2) / INPUT_SIZE * orig_h)
    # 裁剪越界坐标
    x1, y1 = max(0, x1), max(0, y1)
    x2, y2 = min(orig_w, x2), min(orig_h, y2)
    candidate_boxes.append([x1, y1, x2, y2])
    candidate_scores.append(float(score))

# NMS去重
final_results = []
if candidate_boxes:
    candidate_boxes = np.array(candidate_boxes)
    candidate_scores = np.array(candidate_scores)
    keep_idx = cv2.dnn.NMSBoxes(
        candidate_boxes.tolist(), 
        candidate_scores.tolist(), 
        CONF_THRESHOLD, 
        IOU_THRESHOLD
    ).flatten()
    final_results = [
        {"box": candidate_boxes[i], "score": candidate_scores[i]}
        for i in keep_idx
    ]
    print(f"检测到{len(final_results)}个'+'字符")
else:
    print("未检测到目标")

注意:如果运行后发现坐标偏移,可检查坐标是否为640尺寸下的绝对像素值(坐标范围0~640),如果是则不需要除以INPUT_SIZE,直接按比例映射回原图即可。官方导出脚本默认会把坐标解码逻辑合并到模型中,无需额外实现Sigmoid、锚框换算步骤。

内容的提问来源于stack exchange,提问作者Maxime R.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 11:06:20