如何全量化Keras H5模型为输入输出均为int8的TFLite模型适配NPU
全int8量化Keras VGGFace模型(TF1.15.3适配)
针对你遇到的「设置了inference_input_type和inference_output_type为tf.int8,但模型输入仍为float32带转换层」的问题,结合社区用户dtlam分享的适配方案,在TF1.15.3下可以通过调整转换器的关键参数实现全int8量化(输入输出均为int8,无额外转换层),核心是启用新转换逻辑并明确指定支持的类型。
关键参数说明
converter.experimental_new_converter = True:在TF1.15.x版本中,启用新的模型转换逻辑,能更准确地响应全量化的输入输出配置,避免残留浮点转换层converter.target_spec.supported_types = [tf.int8]:强制转换器仅支持int8类型,确保从输入到输出的所有层都采用int8量化,而非仅中间层representative_dataset_gen:必须提供符合模型输入格式的浮点样本数据,用于校准量化过程中的数值分布,这是全量化的必要前提
完整量化代码
import tensorflow as tf import cv2 import numpy as np saved_model_dir = "你的模型目录路径/" modelname = "keras_vggface.h5" converter = tf.lite.TFLiteConverter.from_keras_model_file(saved_model_dir + modelname) def representative_dataset_gen(): # 注意:这里需要提供至少10张符合模型输入要求的样本图片 for _ in range(10): pfad='pathtoimage/000001.jpg' # 替换为你的样本图片路径 img=cv2.imread(pfad) # 确保图片预处理逻辑和你训练/推理时一致,比如尺寸、归一化等 img = np.expand_dims(img, 0).astype(np.float32) yield [img] converter.representative_dataset = representative_dataset_gen converter.optimizations = [tf.lite.Optimize.DEFAULT] converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8] # 关键新增参数 converter.experimental_new_converter = True converter.target_spec.supported_types = [tf.int8] converter.inference_input_type = tf.int8 converter.inference_output_type = tf.int8 quantized_tflite_model = converter.convert() # 保存量化后的模型 if tf.__version__.startswith('1.'): open("test153.tflite", "wb").write(quantized_tflite_model)
验证方法
量化完成后,可以用Netron等模型可视化工具打开生成的tflite文件,检查:
- 输入层的类型为
int8,没有额外的Quantize/Dequantize转换层 - 所有中间层的输入、输出、权重均为int8(偏置通常为int32,这是正常现象)
- 输出层类型为
int8
内容的提问来源于stack exchange,提问作者Florida Man
相关产品推荐
相关产品推荐

