You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TFLite量化Dense层计算逻辑复现异常排查求助

量化Dense层计算复现与TFLite输出不符问题

我正在尝试明确TFLite中量化Dense层的具体计算逻辑,参考量化原理相关内容做代码复现,但使用从TFLite模型提取的参数计算时,结果和模型实际输出不一致,所有值出现饱和现象。

以下是我的复现代码:

import tensorflow as tf
from tensorflow import keras
import numpy as np

i = 16
j = 48
k = 32

model = tf.keras.models.Sequential()
model.add(tf.keras.Input(shape=(i, j), batch_size=1))
model.add(tf.keras.layers.Dense(k, activation=None, use_bias=False))

def representative_data_gen():
    dataset = [
        np.array(
            np.random.randint(-10, 10, size=(i, j)),
            dtype=np.float32,
        )
        for s in range(10)
    ]
    for input_value in dataset:
        # Model has only one input so each data point has one element.s
        yield [input_value] 

converter = tf.lite.TFLiteConverter.from_keras_model(model)

converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_data_gen
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS, tf.lite.OpsSet.SELECT_TF_OPS, tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.target_spec.supported_types = [tf.int8]
converter.inference_input_type = tf.int8
converter.inference_output_type = tf.int8
tflite_model_quant = converter.convert()

# Run the model with TensorFlow Lite
interpreter = tf.lite.Interpreter(model_content=tflite_model_quant)
interpreter.allocate_tensors()

tensor_details = interpreter.get_tensor_details()
out_details = interpreter.get_output_details()
input_tensor = np.squeeze(interpreter.get_tensor(input_idx)) #shape: [i, j]
weight_tensor = np.squeeze(interpreter.get_tensor(weight_idx)) #shape: [k, j]
out_tensor = np.squeeze(interpreter.get_tensor(out_idx))

reduction_dim_size = weight_tensor.shape[1]

zp_input = tensor_details[input_idx]["quantization_parameters"]["zero_points"][0]
q_input = tensor_details[input_idx]["quantization_parameters"]["scales"][0]

zp_weight = tensor_details[weight_idx]["quantization_parameters"]["zero_points"][0]
q_weight = tensor_details[weight_idx]["quantization_parameters"]["scales"][0]

zp_out = tensor_details[out_idx]["quantization_parameters"]["zero_points"][0]
q_out = tensor_details[out_idx]["quantization_parameters"]["scales"][0]

result = (zp_out +
    ((q_input*q_weight)/q_out) +
    np.matmul(input_tensor, weight_tensor.T)
    - zp_weight * np.sum(input_tensor.astype(np.int32), axis=1, keepdims=True)
    - zp_input * np.sum(weight_tensor.T.astype(np.int32), axis=0, keepdims=True)
    + reduction_dim_size*zp_input*zp_weight)
result = np.round(result, decimals=0)
result = np.clip(result, a_min=-128, a_max=127)
result = result.astype(np.int8)

我怀疑计算式里zp_input * np.sum(weight_tensor.T.astype(np.int32), axis=0, keepdims=True)这一项存在问题,但无法定位具体原因。此外,我找到了TFLite官方的整数运算Dense层实现代码,对其中的输出移位值来源存疑,推测它与量化scale因子相关,但不确定具体的对应关系。

内容的提问来源于stack exchange,提问作者Necrotos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 02:46:17