如何从Qualcomm SNPE量化YOLOv8的Hexagon DSP输出绘制检测框
解决SNPE量化YOLOv8s后反量化输出异常及检测框绘制问题
1. 修正反量化公式(核心问题)
你当前使用的反量化公式顺序错误,SNPE INT8量化的正确反量化逻辑是:
dequant_output = (quant_output - OFFSET) * SCALE
而非 quant_output * SCALE - OFFSET。SNPE量化激活层时,是将浮点值先除以缩放因子、取整后加上偏移量得到INT8值;反量化需要逆向操作:先减去偏移量,再乘以缩放因子。顺序错误会直接导致数值被过度放大,造成置信度过高的问题。
2. 正确读取并解析SNPE输出的RAW文件
2.1 确认输出张量结构
先通过snpe-dlc-info -i yolov8s_quantized.dlc查看输出节点output0的详细信息,重点关注:
- 输出维度(二分类YOLOv8s的输出通常为
1,6,8400,对应[batch, 5+num_classes, num_detections]) - 输出数据类型(量化后应为
INT8)
2.2 读取RAW二进制文件
以Python为例,按照INT8格式读取RAW文件后,reshape为对应维度:
import numpy as np # 读取RAW输出文件 raw_output = np.fromfile("output/output0.raw", dtype=np.int8) # 根据dlc-info的输出维度reshape,示例为1,6,8400 output_tensor = raw_output.reshape(1, 6, 8400).transpose(0, 2, 1) # 转换为(1, 8400, 6),每个检测框对应一行 detections = output_tensor[0] # 取batch中的第一个样本
3. 执行反量化并解析检测结果
3.1 反量化处理
使用修正后的公式对检测结果反量化:
SCALE = 2.875693559647 OFFSET = -15 # 对每个检测框的所有字段反量化 dequant_detections = (detections - OFFSET) * SCALE
3.2 解析YOLO输出并绘制检测框
YOLOv8的每个检测框字段为[x_center, y_center, width, height, confidence, class_score],按以下步骤处理:
import cv2 # 原始图像信息 original_img = cv2.imread("sample.jpg") orig_h, orig_w = original_img.shape[:2] input_size = 640 # 置信度阈值和NMS阈值 conf_thresh = 0.5 nms_thresh = 0.5 # 过滤低置信度框 valid_detections = dequant_detections[dequant_detections[:,4] >= conf_thresh] # 坐标转换:将模型输入尺寸的框映射到原始图像尺寸 def scale_coords(input_size, coords, img_shape): gain = min(input_size/img_shape[1], input_size/img_shape[0]) pad = (input_size - img_shape[1]*gain)/2, (input_size - img_shape[0]*gain)/2 coords[:, [0,2]] -= pad[0] coords[:, [1,3]] -= pad[1] coords[:, :4] /= gain # 确保坐标在图像范围内 coords[:, 0] = np.clip(coords[:, 0], 0, img_shape[1]-1) coords[:, 1] = np.clip(coords[:, 1], 0, img_shape[0]-1) coords[:, 2] = np.clip(coords[:, 2], 0, img_shape[1]-1) coords[:, 3] = np.clip(coords[:, 3], 0, img_shape[0]-1) return coords # 将中心坐标转换为左上角/右下角坐标 boxes = valid_detections[:, :4].copy() boxes[:, 0] = valid_detections[:, 0] - valid_detections[:, 2]/2 boxes[:, 1] = valid_detections[:, 1] - valid_detections[:, 3]/2 boxes[:, 2] = valid_detections[:, 0] + valid_detections[:, 2]/2 boxes[:, 3] = valid_detections[:, 1] + valid_detections[:, 3]/2 # 映射到原始图像尺寸 boxes = scale_coords(input_size, boxes, (orig_h, orig_w)) # 执行非极大值抑制(NMS)去重 indices = cv2.dnn.NMSBoxes(boxes.tolist(), valid_detections[:,4].tolist(), conf_thresh, nms_thresh) # 绘制检测框和标签 for i in indices: i = i[0] if isinstance(i, (list, np.ndarray)) else i x1, y1, x2, y2 = map(int, boxes[i]) conf = valid_detections[i,4] cls = valid_detections[i,5] cv2.rectangle(original_img, (x1,y1), (x2,y2), (0,255,0), 2) label = f"Class {int(cls)}: {conf:.2f}" cv2.putText(original_img, label, (x1,y1-10), cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0,255,0), 2) # 保存或显示结果 cv2.imwrite("result.jpg", original_img) cv2.imshow("Detection Result", original_img) cv2.waitKey(0)
4. 额外检查点
- 校准数据一致性:确认量化时
raw_list.txt中的校准图像,与推理输入的图像格式完全一致(NCHW、float32、除以255归一化),校准数据偏差会导致量化模型输出异常。 - DSP量化兼容性:如果使用DSP推理,量化时建议添加
--enable_htp_quantization参数(针对Qualcomm HTP),避免层量化不兼容导致输出错误。
内容的提问来源于stack exchange,提问作者Lap Nguyen
相关产品推荐
相关产品推荐

