You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将YOLOv5 Nano转为8位量化TF-Lite Micro及归一化输出框?

问题1:将YOLOv5 Nano转换为8位量化TF-Lite

由于你的输出张量同时包含0-640的边界框坐标和0-1的置信度,共享量化参数会导致精度异常,核心解决思路是拆分输出张量,让两类数据分别适配量化参数:

  • 步骤1:修改YOLOv5模型输出结构
    打开models/yolo.py,找到Detect类的forward方法,将原本合并的输出拆分为两个独立张量:一个存储边界框坐标,另一个存储置信度/类别概率。示例修改如下:

    def forward(self, x):
        # 保留原前向计算逻辑,得到pred张量
        pred = self.forward_once(x)
        # 拆分张量:前4列为边界框坐标,剩余列为置信度/类别概率
        boxes = pred[..., :4]
        scores = pred[..., 4:]
        return boxes, scores  # 返回两个独立输出
    
  • 步骤2:重新导出TensorFlow SavedModel
    使用Ultralytics官方导出命令,导出为TensorFlow SavedModel格式:

    python export.py --weights yolov5n.pt --include saved_model --img 640
    
  • 步骤3:执行8位量化转换
    使用TensorFlow Lite Python API进行量化,指定代表性数据集校准,让转换器为两个输出分别计算适配的缩放因子和零点:

    import tensorflow as tf
    
    # 加载SavedModel
    converter = tf.lite.TFLiteConverter.from_saved_model("yolov5n_saved_model")
    # 启用默认优化(包含8位量化)
    converter.optimizations = [tf.lite.Optimize.DEFAULT]
    # 提供代表性数据集生成器(替换为你的数据集路径)
    def representative_data_gen():
        for image_path in ["sample1.jpg", "sample2.jpg"]:
            img = tf.io.read_file(image_path)
            img = tf.image.decode_jpeg(img, channels=3)
            img = tf.image.resize(img, (640, 640))
            img = tf.cast(img, tf.float32) / 255.0
            yield [tf.expand_dims(img, 0)]
    converter.representative_dataset = representative_data_gen
    # 设置输入输出类型为uint8(8位量化)
    converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
    converter.inference_input_type = tf.uint8
    converter.inference_output_type = tf.uint8
    
    # 生成量化后的TF-Lite模型
    tflite_model = converter.convert()
    with open("yolov5n_quantized.tflite", "wb") as f:
        f.write(tflite_model)
    
问题2:导出Ultralytics模型使边界框输出为0-1归一化值

直接在模型前向传播阶段对边界框坐标做归一化处理,再导出即可:

  • 方法1:修改YOLOv5模型输出逻辑
    在models/yolo.py的Detect类forward方法中,对边界框坐标除以输入图像的宽高:

    def forward(self, x):
        pred = self.forward_once(x)
        # 动态获取输入图像的宽高
        img_h, img_w = x.shape[2], x.shape[3]
        # 对边界框坐标做归一化
        pred[..., 0] /= img_w  # x坐标归一化
        pred[..., 1] /= img_h  # y坐标归一化
        pred[..., 2] /= img_w  # 宽度归一化
        pred[..., 3] /= img_h  # 高度归一化
        return pred
    

    修改完成后,用常规导出命令导出TF-Lite或其他格式,输出的边界框坐标就会是0-1的归一化值。

  • 方法2:导出后添加后处理层(TF-Lite专属)
    如果不想修改原模型,可以在导出SavedModel后,添加自定义后处理层完成归一化,再进行量化:

    import tensorflow as tf
    
    # 加载原SavedModel
    loaded_model = tf.saved_model.load("yolov5n_saved_model")
    infer_func = loaded_model.signatures["serving_default"]
    
    # 定义带归一化的新模型
    class NormalizedYOLO(tf.keras.Model):
        def __init__(self, original_model):
            super().__init__()
            self.original_model = original_model
    
        @tf.function(input_signature=[tf.TensorSpec(shape=[None, 640, 640, 3], dtype=tf.float32)])
        def call(self, inputs):
            outputs = self.original_model(inputs)
            pred = outputs["output_0"]
            # 归一化边界框(假设输入尺寸为640x640)
            pred = tf.concat([
                pred[..., :4] / 640.0,
                pred[..., 4:]
            ], axis=-1)
            return pred
    
    # 保存新模型
    new_model = NormalizedYOLO(loaded_model)
    tf.saved_model.save(new_model, "yolov5n_normalized_saved_model")
    
    # 后续按常规流程量化为TF-Lite即可
    

内容的提问来源于stack exchange,提问作者PolarBear2015

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 02:47:22