YOLOv8导出为ONNX后性能劣于.pt模型,求GPU环境下解决方法
YOLOv8 ONNX导出后性能对齐解决方案
核心问题排查与解决步骤
矩形推理参数严格对齐
训练时启用了rect=True和1280×720的矩形输入,导出与推理阶段必须完全匹配该逻辑:- 导出ONNX时,
imgsz参数顺序需与训练一致(应为1280,720,对应YOLOv8的高、宽定义),且必须加上rect=True参数,避免导出为正方形输入模型。 - 推理阶段必须复现YOLOv8的矩形预处理逻辑:按图像宽高比缩放,仅对短边做均值(默认114)填充,不能用常规正方形padding。
- 导出ONNX时,
导出参数精简匹配
- 暂时关闭
simplify参数,部分场景下模型简化会修改算子结构导致精度损失,使用以下命令导出:yolo task=detect mode=export model=runs/detect/last.pt imgsz=1280,720 rect=True format=onnx opset=12 device=0 - 指定
device=0确保用GPU导出,避免CPU导出带来的算子差异。
- 暂时关闭
ONNX Runtime推理细节对齐
- 推理时指定GPU执行器,设置
providers=['CUDAExecutionProvider', 'CPUExecutionProvider'],强制用GPU加速推理。 - 预处理严格匹配YOLOv8逻辑:图像转RGB、除以255.0归一化到0-1区间,无需额外减均值。
- 后处理复现YOLOv8的NMS规则:使用相同的置信度阈值、IOU阈值,同时将检测框坐标从预处理后的画布尺寸还原回原图像尺寸。
- 推理时指定GPU执行器,设置
模型精度验证对比
- 先用PT模型跑验证:
yolo task=detect mode=val model=runs/detect/last.pt data=your_data.yaml rect=True - 再用ONNX模型跑验证:
yolo task=detect mode=val model=runs/detect/last.onnx data=your_data.yaml rect=True - 对比两者mAP指标,若差距大说明导出环节有问题;若差距小则是推理代码的预处理/后处理未对齐。
- 先用PT模型跑验证:
GPU环境下亲测可行的推理流程
- 导出正确的ONNX模型
yolo task=detect mode=export model=runs/detect/last.pt imgsz=1280,720 rect=True format=onnx opset=12 device=0
- ONNX Runtime推理核心代码片段
import onnxruntime as ort import cv2 import numpy as np # 加载模型并指定GPU执行器 sess = ort.InferenceSession("last.onnx", providers=['CUDAExecutionProvider']) # 矩形推理预处理 def preprocess(img, target_size=(1280, 720)): h, w = img.shape[:2] scale = min(target_size[0]/h, target_size[1]/w) new_h, new_w = int(h*scale), int(w*scale) img_resized = cv2.resize(img, (new_w, new_h), interpolation=cv2.INTER_LINEAR) # 用均值114填充画布 canvas = np.full((target_size[0], target_size[1], 3), 114, dtype=np.uint8) canvas[:new_h, :new_w] = img_resized # 转RGB、归一化、调整维度为NCHW img_rgb = cv2.cvtColor(canvas, cv2.COLOR_BGR2RGB) img_tensor = img_rgb.astype(np.float32) / 255.0 img_tensor = np.transpose(img_tensor, (2, 0, 1))[None, :, :, :] return img_tensor, scale, (new_w, new_h) # 后处理还原框坐标并执行NMS def postprocess(outputs, scale, new_shape, img_shape, conf_thres=0.25, iou_thres=0.7): predictions = outputs[0] boxes = predictions[:, :4] confs = predictions[:, 4:5] * predictions[:, 5:] class_ids = np.argmax(confs, axis=1) confs = np.max(confs, axis=1) # 过滤低置信度框 mask = confs > conf_thres boxes, confs, class_ids = boxes[mask], confs[mask], class_ids[mask] # 还原框到原图像尺寸 boxes[:, [0, 2]] = (boxes[:, [0, 2]] - (1280 - new_shape[0])/2) / scale boxes[:, [1, 3]] = (boxes[:, [1, 3]] - (720 - new_shape[1])/2) / scale # 限制框在图像范围内 boxes = np.clip(boxes, 0, [img_shape[1], img_shape[0], img_shape[1], img_shape[0]]) # 执行NMS indices = cv2.dnn.NMSBoxes(boxes.tolist(), confs.tolist(), conf_thres, iou_thres) if len(indices) > 0: indices = indices.flatten() return boxes[indices], confs[indices], class_ids[indices] return [], [], [] # 推理示例 img = cv2.imread("test.jpg") img_tensor, scale, new_shape = preprocess(img) outputs = sess.run(None, {sess.get_inputs()[0].name: img_tensor}) boxes, confs, class_ids = postprocess(outputs, scale, new_shape, img.shape[:2])
内容的提问来源于stack exchange,提问作者moonboi
相关产品推荐
相关产品推荐

