You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何全量化Keras H5模型为输入输出均为int8的TFLite模型适配NPU

全int8量化Keras VGGFace模型(TF1.15.3适配)

针对你遇到的「设置了inference_input_type和inference_output_type为tf.int8,但模型输入仍为float32带转换层」的问题,结合社区用户dtlam分享的适配方案,在TF1.15.3下可以通过调整转换器的关键参数实现全int8量化(输入输出均为int8,无额外转换层),核心是启用新转换逻辑并明确指定支持的类型。

关键参数说明

  • converter.experimental_new_converter = True:在TF1.15.x版本中,启用新的模型转换逻辑,能更准确地响应全量化的输入输出配置,避免残留浮点转换层
  • converter.target_spec.supported_types = [tf.int8]:强制转换器仅支持int8类型,确保从输入到输出的所有层都采用int8量化,而非仅中间层
  • representative_dataset_gen:必须提供符合模型输入格式的浮点样本数据,用于校准量化过程中的数值分布,这是全量化的必要前提

完整量化代码

import tensorflow as tf
import cv2
import numpy as np

saved_model_dir = "你的模型目录路径/"
modelname = "keras_vggface.h5"

converter = tf.lite.TFLiteConverter.from_keras_model_file(saved_model_dir + modelname)

def representative_dataset_gen():
    # 注意:这里需要提供至少10张符合模型输入要求的样本图片
    for _ in range(10):
        pfad='pathtoimage/000001.jpg'  # 替换为你的样本图片路径
        img=cv2.imread(pfad)
        # 确保图片预处理逻辑和你训练/推理时一致,比如尺寸、归一化等
        img = np.expand_dims(img, 0).astype(np.float32)
        yield [img]

converter.representative_dataset = representative_dataset_gen
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
# 关键新增参数
converter.experimental_new_converter = True
converter.target_spec.supported_types = [tf.int8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8

quantized_tflite_model = converter.convert()

# 保存量化后的模型
if tf.__version__.startswith('1.'):
    open("test153.tflite", "wb").write(quantized_tflite_model)

验证方法

量化完成后,可以用Netron等模型可视化工具打开生成的tflite文件,检查:

  • 输入层的类型为int8,没有额外的Quantize/Dequantize转换层
  • 所有中间层的输入、输出、权重均为int8(偏置通常为int32,这是正常现象)
  • 输出层类型为int8

内容的提问来源于stack exchange,提问作者Florida Man

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 08:27:47