You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Qualcomm SNPE量化YOLOv8的Hexagon DSP输出绘制检测框

解决SNPE量化YOLOv8s后反量化输出异常及检测框绘制问题

1. 修正反量化公式(核心问题)

你当前使用的反量化公式顺序错误,SNPE INT8量化的正确反量化逻辑是:

dequant_output = (quant_output - OFFSET) * SCALE

而非 quant_output * SCALE - OFFSET。SNPE量化激活层时,是将浮点值先除以缩放因子、取整后加上偏移量得到INT8值;反量化需要逆向操作:先减去偏移量,再乘以缩放因子。顺序错误会直接导致数值被过度放大,造成置信度过高的问题。

2. 正确读取并解析SNPE输出的RAW文件

2.1 确认输出张量结构

先通过snpe-dlc-info -i yolov8s_quantized.dlc查看输出节点output0的详细信息,重点关注:

  • 输出维度(二分类YOLOv8s的输出通常为1,6,8400,对应[batch, 5+num_classes, num_detections])
  • 输出数据类型(量化后应为INT8)

2.2 读取RAW二进制文件

以Python为例,按照INT8格式读取RAW文件后,reshape为对应维度:

import numpy as np

# 读取RAW输出文件
raw_output = np.fromfile("output/output0.raw", dtype=np.int8)
# 根据dlc-info的输出维度reshape,示例为1,6,8400
output_tensor = raw_output.reshape(1, 6, 8400).transpose(0, 2, 1)  # 转换为(1, 8400, 6),每个检测框对应一行
detections = output_tensor[0]  # 取batch中的第一个样本

3. 执行反量化并解析检测结果

3.1 反量化处理

使用修正后的公式对检测结果反量化:

SCALE = 2.875693559647
OFFSET = -15

# 对每个检测框的所有字段反量化
dequant_detections = (detections - OFFSET) * SCALE

3.2 解析YOLO输出并绘制检测框

YOLOv8的每个检测框字段为[x_center, y_center, width, height, confidence, class_score],按以下步骤处理:

import cv2

# 原始图像信息
original_img = cv2.imread("sample.jpg")
orig_h, orig_w = original_img.shape[:2]
input_size = 640

# 置信度阈值和NMS阈值
conf_thresh = 0.5
nms_thresh = 0.5

# 过滤低置信度框
valid_detections = dequant_detections[dequant_detections[:,4] >= conf_thresh]

# 坐标转换:将模型输入尺寸的框映射到原始图像尺寸
def scale_coords(input_size, coords, img_shape):
    gain = min(input_size/img_shape[1], input_size/img_shape[0])
    pad = (input_size - img_shape[1]*gain)/2, (input_size - img_shape[0]*gain)/2
    coords[:, [0,2]] -= pad[0]
    coords[:, [1,3]] -= pad[1]
    coords[:, :4] /= gain
    # 确保坐标在图像范围内
    coords[:, 0] = np.clip(coords[:, 0], 0, img_shape[1]-1)
    coords[:, 1] = np.clip(coords[:, 1], 0, img_shape[0]-1)
    coords[:, 2] = np.clip(coords[:, 2], 0, img_shape[1]-1)
    coords[:, 3] = np.clip(coords[:, 3], 0, img_shape[0]-1)
    return coords

# 将中心坐标转换为左上角/右下角坐标
boxes = valid_detections[:, :4].copy()
boxes[:, 0] = valid_detections[:, 0] - valid_detections[:, 2]/2
boxes[:, 1] = valid_detections[:, 1] - valid_detections[:, 3]/2
boxes[:, 2] = valid_detections[:, 0] + valid_detections[:, 2]/2
boxes[:, 3] = valid_detections[:, 1] + valid_detections[:, 3]/2

# 映射到原始图像尺寸
boxes = scale_coords(input_size, boxes, (orig_h, orig_w))

# 执行非极大值抑制(NMS)去重
indices = cv2.dnn.NMSBoxes(boxes.tolist(), valid_detections[:,4].tolist(), conf_thresh, nms_thresh)

# 绘制检测框和标签
for i in indices:
    i = i[0] if isinstance(i, (list, np.ndarray)) else i
    x1, y1, x2, y2 = map(int, boxes[i])
    conf = valid_detections[i,4]
    cls = valid_detections[i,5]
    cv2.rectangle(original_img, (x1,y1), (x2,y2), (0,255,0), 2)
    label = f"Class {int(cls)}: {conf:.2f}"
    cv2.putText(original_img, label, (x1,y1-10), cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0,255,0), 2)

# 保存或显示结果
cv2.imwrite("result.jpg", original_img)
cv2.imshow("Detection Result", original_img)
cv2.waitKey(0)

4. 额外检查点

  • 校准数据一致性:确认量化时raw_list.txt中的校准图像,与推理输入的图像格式完全一致(NCHW、float32、除以255归一化),校准数据偏差会导致量化模型输出异常。
  • DSP量化兼容性:如果使用DSP推理,量化时建议添加--enable_htp_quantization参数(针对Qualcomm HTP),避免层量化不兼容导致输出错误。

内容的提问来源于stack exchange,提问作者Lap Nguyen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 14:39:55