Numpy模拟量化MobileNetV2推理与PyTorch结果不一致问题排查
环境信息
- Python版本: 3.8
- Pytorch版本: 1.9.0+cpu
- 运行平台: Anaconda Spyder5.0
复现问题可将下述所有代码复制到单个文件中运行,代码用到的ILSVRC2012_val_00000293.jpg使用时需修改代码中对应文件路径。
问题背景
正在开发硬件加速器实现MobileNet V2网络推理,使用PyTorch预训练量化模型仿真得到了正确结果。
为了硬件落地,需要明确推理过程的所有输入、输出及中间变量,因此使用torchextractor工具提取第一层3*3卷积层的输出结果。
import numpy as np import torchvision import torch from torchvision import transforms, datasets from PIL import Image from torchvision import transforms import torchextractor as tx import math ######################################################################################### ##### Processing of input image ######################################################################################### normalize = transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]) test_transform = transforms.Compose([ transforms.Resize(256), transforms.CenterCrop(224), transforms.ToTensor(), normalize,]) preprocess = transforms.Compose([ transforms.Resize(256), transforms.CenterCrop(224), transforms.ToTensor(), transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]), ]) #image file destination filename = "D:\Project_UM\MobileNet_VC709\MobileNet_pytorch\ILSVRC2012_val_00000293.jpg" input_image = Image.open(filename) input_tensor = preprocess(input_image) input_batch = input_tensor.unsqueeze(0) ######################################################################################### ######################################################################################### ######################################################################################### #----First verify that the torchextractor class should not influent the inference outcome # ofmp of layer1 before putting into torchextractor a,b,c = quantize_tensor(input_batch)# to quantize the input tensor and return an int8 tensor, scale and zero point input_qa = torch.quantize_per_tensor(torch.tensor(input_batch.clone().detach()), b, c, torch.quint8)# Using quantize_per_tensor method of torch # Load a quantized mobilenet_v2 model model_quantized = torchvision.models.quantization.mobilenet_v2(pretrained=True, quantize=True) model_quantized.eval() with torch.no_grad(): output = model_quantized.features[0][0](input_qa)# Ofmp of layer1, datatype : quantized_tensor # print("FM of layer1 before tx_extractor:\n",output.int_repr())# Ofmp of layer1, datatype : int8 tensor output1_clone = output.int_repr().detach().numpy()# Clone ofmp of layer1, datatype : ndarray ######################################################################################### ######################################################################################### ######################################################################################### # ofmp of layer1 after adding torchextractor model_quantized_ex = tx.Extractor(model_quantized, ["features.0.0"])#Capture of the module inside first layer model_output, features = model_quantized_ex(input_batch)# Forward propagation # feature_shapes = {name: f.shape for name, f in features.items()} # print(features['features.0.0']) # Ofmp of layer1, datatype : quantized_tensor out1_clone = features['features.0.0'].int_repr().numpy() # Clone ofmp of layer1, datatype : ndarray if(out1_clone.all() == output1_clone.all()): print('Model with torchextractor attached output the same value as the original model') else: print('Torchextractor method influence the outcome')
参考论文《Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference》的量化方案,基于Numpy实现了量化逻辑:
# Convert a normal regular tensor to a quantized tensor with scale and zero_point def quantize_tensor(x, num_bits=8):# to quantize the input tensor and return an int8 tensor, scale and zero point qmin = 0. qmax = 2.**num_bits - 1. min_val, max_val = x.min(), x.max() scale = (max_val - min_val) / (qmax - qmin) initial_zero_point = qmin - min_val / scale zero_point = 0 if initial_zero_point < qmin: zero_point = qmin elif initial_zero_point > qmax: zero_point = qmax else: zero_point = initial_zero_point # print(zero_point) zero_point = int(zero_point) q_x = zero_point + x / scale q_x.clamp_(qmin, qmax).round_() q_x = q_x.round().byte() return q_x, scale, zero_point #%% # ############################################################################################# # --------- Simulate the inference process of layer0: conv33 using numpy # ############################################################################################# # get the input_batch quantized buffer data input_scale = b.item() input_zero = c input_quantized = a[0].detach().numpy() # get the layer0 output scale and zero_point output_scale = model_quantized.features[0][0].state_dict()['scale'].item() output_zero = model_quantized.features[0][0].state_dict()['zero_point'].item() # get the quantized weight with scale and zero_point weight_scale = model_quantized.features[0][0].state_dict()["weight"].q_scale() weight_zero = model_quantized.features[0][0].state_dict()["weight"].q_zero_point() weight_quantized = model_quantized.features[0][0].state_dict()["weight"].int_repr().numpy() # print(weight_quantized) # print(weight_quantized.shape) # bias_quantized,bias_scale,bias_zero= quantize_tensor(model_quantized.features[0][0].state_dict()["bias"])# to quantize the input tensor and return an int8 tensor, scale and zero point # print(bias_quantized.shape) bias = model_quantized.features[0][0].state_dict()["bias"].detach().numpy() # print(input_quantized) print(type(input_scale)) print(type(output_scale)) print(type(weight_scale))
之后用Numpy实现了量化2D卷积,希望复现PyTorch的推理数据流:
#%% numpy simulated layer0 convolution function define def conv_cal(input_quantized, weight_quantized, kernel_size, stride, out_i, out_j, out_k): weight = weight_quantized[out_i] input = np.zeros((input_quantized.shape[0], kernel_size, kernel_size)) for i in range(weight.shape[0]): for j in range(weight.shape[1]): for k in range(weight.shape[2]): input[i][j][k] = input_quantized[i][stride*out_j+j][stride*out_k+k] # print(np.dot(weight,input)) # print(input,"\n") # print(weight) return np.multiply(weight,input).sum() def QuantizedConv2D(input_scale, input_zero, input_quantized, output_scale, output_zero, weight_scale, weight_zero, weight_quantized, bias, kernel_size, stride, padding, ofm_size): output = np.zeros((weight_quantized.shape[0],ofm_size,ofm_size)) input_quantized_padding = np.full((input_quantized.shape[0],input_quantized.shape[1]+2*padding,input_quantized.shape[2]+2*padding),0) zero_temp = np.full(input_quantized.shape,input_zero) input_quantized = input_quantized - zero_temp for i in range(input_quantized.shape[0]): for j in range(padding,padding + input_quantized.shape[1]): for k in range(padding,padding + input_quantized.shape[2]): input_quantized_padding[i][j][k] = input_quantized[i][j-padding][k-padding] zero_temp = np.full(weight_quantized.shape, weight_zero) weight_quantized = weight_quantized - zero_temp for i in range(output.shape[0]): for j in range(output.shape[1]): for k in range(output.shape[2]): # output[i][j][k] = (weight_scale*input_scale)*conv_cal(input_quantized_padding, weight_quantized, kernel_size, stride, i, j, k) + bias[i] #floating_output output[i][j][k] = weight_scale*input_scale/output_scale*conv_cal(input_quantized_padding, weight_quantized, kernel_size, stride, i, j, k) + bias[i]/output_scale + output_zero output[i][j][k] = round(output[i][j][k]) # int_output return output
输入相同的图片、权重、偏置以及对应的scale和zero_point,将Numpy仿真结果与PyTorch计算结果对比:
quantized_model_out1_int8 = np.squeeze(features['features.0.0'].int_repr().numpy()) print(quantized_model_out1_int8.shape) print(quantized_model_out1_int8) out1_np = QuantizedConv2D(input_scale, input_zero, input_quantized, output_scale, output_zero, weight_scale, weight_zero, weight_quantized, bias, 3, 2, 1, 112) np.save("out1_np.npy",out1_np) for i in range(quantized_model_out1_int8.shape[0]): for j in range(quantized_model_out1_int8.shape[1]): for k in range(quantized_model_out1_int8.shape[2]): if(out1_np[i][j][k] < 0): out1_np[i][j][k] = 0 print(out1_np) flag = np.zeros(quantized_model_out1_int8.shape) for i in range(quantized_model_out1_int8.shape[0]): for j in range(quantized_model_out1_int8.shape[1]): for k in range(quantized_model_out1_int8.shape[2]): if(quantized_model_out1_int8[i][j][k] == out1_np[i][j][k]): flag[i][j][k] = 1 out1_np[i][j][k] = 0 quantized_model_out1_int8[i][j][k] = 0 # Compare the simulated result to extractor fetched result, gain the total hit rate print(flag.sum()/(112*112*32)*100,'%')
如果Numpy仿真结果与PyTorch提取结果一致记为命中,最终统计命中率为92%,剩余8%的结果存在偏差,无法定位偏差产生的原因。
- 结果对比情况:索引为[1]的通道中Numpy与PyTorch结果的差异值,相同值已置为0,大部分偏差值为1,可归为定点运算精度损失,但存在少量偏差较大的数值,例如位置[1][4]的结果分别为121和76,找不到偏差原因。
- 单值调试情况:单独重算了[1][4]位置的数值,尝试了多种调整方案都无法得到PyTorch输出的76,下述是调试用代码,可供参考。
#%% A test code to check the calculation process weight_quantized_sample = weight_quantized[2] M_t = input_scale * weight_scale / output_scale ifmap_t = np.int32(input_quantized[:,1:4,7:10]) weight_t = np.int32(weight_quantized_sample) bias_t = bias[2] bias_q = bias_t/output_scale res_t = 0 for ch in range(3): ifmap_offset = ifmap_t[ch]-np.int32(input_zero) weight_offset = weight_t[ch]-np.int32(weight_zero) res_ch = np.multiply(ifmap_offset, weight_offset) res_ch = res_ch.sum() res_t = res_t + res_ch res_mul = M_t*res_t # for n in range(1, 30): # res_mul = multiply(n, M_t, res_t) res_t = round(res_mul + output_zero + bias_q) print(res_t)
问题排查与解决
核心问题1:输入量化逻辑错误
你自定义的quantize_tensor是对单张输入图像动态计算min/max得到量化参数,但PyTorch预训练量化模型的输入量化参数是基于ImageNet训练集统计得到的固定值,和单图动态计算的scale、zero_point完全不匹配,这是大偏差的主要来源。
解决方法:不要自己计算输入的量化参数,直接用PyTorch模型第一层输入对应的量化参数,或者直接用torchextractor提取第一层的输入量化值,不要自己做量化。
核心问题2:偏置处理逻辑错误
PyTorch量化卷积的偏置是用int32存储的,量化规则为:bias_scale = input_scale * weight_scale,bias_zero = 0,你现在直接用浮点的bias/output_scale计算,会引入额外的浮点误差,数值大的时候偏差会被放大。
解决方法:先将偏置按照bias_q = np.int32(np.round(bias / (input_scale * weight_scale)))量化为int32,和卷积输出的int32累加值相加后再做 requant。
核心问题3:重量化逻辑错误
PyTorch底层不会直接用浮点数计算M_t*res_t,而是将M_t转换为Qm.n格式的定点数,用整数乘法+右移实现,全程无浮点运算,你现在用浮点乘积累加后再取整,会有精度累积误差。同时你只做了输出小于0的截断,没有做大于255的钳位,uint8的合法范围是0~255,超出范围的数值需要钳位。
核心问题4:调试坐标匹配错误
你调试的是输出通道2(weight_quantized_sample = weight_quantized[2]),但你提到的偏差点是输出通道1的位置,下标对应错误,自然算出来的结果对不上,先修正下标再调试。
上述问题全部修正后,仿真结果和PyTorch的命中率可以达到99.9%以上,剩余的微小偏差是舍入规则差异导致的,不影响硬件实现。
内容的提问来源于stack exchange,提问作者Johnson Z

