You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

YOLOv11n自定义推理与CLI预测结果不符问题求助

问题描述

我使用微调后导出的YOLOv11n模型检测路标,通过YOLO CLI对同一张图片推理时,能得到置信度为0.86的目标检测结果;但用自定义ONNX Runtime脚本推理时,所有置信度均在1e-7左右,无有效检测结果。我已尝试TensorFlow、Pillow、OpenCV及多种归一化预处理方法,均未解决问题。模型本身可用,问题应出在预处理或后处理环节,求排查思路。

(附Netron架构图,展示模型输入/输出结构)

排查思路
  • 预处理环节对齐YOLO官方流程

    • 图像通道顺序:YOLO默认使用RGB格式,但OpenCV读取图像为BGR格式,代码中未做通道转换,这是最可能的核心问题,需添加img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)步骤。
    • 归一化逻辑:确认YOLOv11导出时的预处理规则,部分版本采用(img / 255.0 - mean) / std而非直接除以255,需与导出参数一致。可查看导出命令中的normalize配置。
    • Resize插值方式:YOLO官方使用双线性插值,需显式指定cv2.resize的插值参数为cv2.INTER_LINEAR,避免默认插值方式差异。
    • 输入张量校验:打印预处理后的张量形状、数据类型及数值范围,确保与模型输入要求(如float32、0-1区间)完全匹配。
  • 后处理环节匹配YOLOv11输出格式

    • 输出维度校验:打印ONNX模型输出的output[0].shape,确认是否为[1, 8400, 4 + num_classes]或[1, 4 + num_classes, 8400],避免转置操作错误。
    • 置信度计算逻辑:YOLOv11的输出通常包含目标置信度(obj_conf)和类别置信度,最终置信度应为obj_conf * 类别置信度。若模型输出格式为[cx, cy, w, h, obj_conf, cls1, cls2, ...],则代码中直接取detection[4:]作为类别得分会遗漏obj_conf,需修正为:
      obj_conf = detection[4]
      class_scores = detection[5:]
      confidence = obj_conf * class_scores[class_id]
      
    • Box解码规则:确认模型输出的box是相对坐标(0-1)还是绝对坐标(相对于640x640),若为相对坐标,需乘以输入尺寸转换为绝对坐标后再计算边界框。
  • ONNX模型与推理环境校验

    • 导出命令一致性:确保使用官方yolo export命令导出ONNX,例如yolo export model=best.pt format=onnx imgsz=640,保证输入尺寸与推理时一致。
    • 输入张量对比:使用YOLO Python API获取推理时的输入张量,在自定义脚本中直接复用该张量进行推理,若输出正常则问题出在预处理;若输出仍异常则需检查ONNX模型导出是否存在问题。
    • ONNX Runtime配置:确认推理会话是否启用了正确的计算后端(如CUDA),避免因精度差异导致结果异常。
自定义ONNX Runtime推理代码
import onnxruntime as ort
import numpy as np
import cv2


class YoloInference:
    # List of class names for YOLO model
    classes = [
    'speed20', 'speed30', 'speed50', 'speed60', 'speed70', 'speed80', 'speed100', 'speed120'
    ]
    
    def __init__(self, model_path):
        self.session = ort.InferenceSession(model_path)
        self.input_name = self.session.get_inputs()[0].name
        self.output_name = self.session.get_outputs()[0].name

    def preprocess(self, image):
        """This functions preprocess the image before inference
        preprocesses:resize(640x640),normalize,HWC to CHW,adds batch
        Returns:
            img preprocessed
        """
        img = cv2.resize(image, (640, 640), interpolation=cv2.INTER_LINEAR)  # 指定双线性插值
        # 新增通道转换:OpenCV读入为BGR,YOLO需要RGB
        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
        img = img.astype(np.float32)
        img = img / 255.0  # Normalize
        img = np.transpose(img, [2, 0, 1])  # Change data layout from HWC to CHW
        img = np.expand_dims(img, axis=0)  # Add batch dimension
        return img

    def postprocess(self, output, conf_threshold=0.5, nms_threshold=0.1):
        detections = output[0].T  # Transpose the array

        boxes = []
        confidences = []
        class_ids = []

        for detection in detections:  # Iterate through each of the 8400 predictions
            box = detection[:4]
            # 修正:提取目标置信度和类别得分
            obj_conf = detection[4]
            class_scores = detection[5:]

            class_id = np.argmax(class_scores)
            # 修正:计算最终置信度
            confidence = obj_conf * class_scores[class_id]
            if confidence > conf_threshold:
                # Convert from center coordinates to corner coordinates
                cx, cy, box_width, box_height = box
                # 若模型输出为相对坐标,转换为绝对坐标
                cx = cx.item() * 640
                cy = cy.item() * 640
                box_height = box_height.item() * 640
                box_width = box_width.item() * 640
                x1 = int((cx - box_width / 2))
                y1 = int((cy - box_height / 2))
                x2 = int(cx + box_width / 2)
                y2 = int(cy + box_height / 2)
                boxes.append([x1, y1, x2, y2])
                confidences.append(float(confidence.item()))
                class_ids.append(class_id)

        # Apply non-maximum suppression (NMS)
        indices = cv2.dnn.NMSBoxes(boxes, confidences, conf_threshold, nms_threshold)
        
        print(confidences)
        
        processed_results = []
        if len(indices) > 0:
            for i in indices.flatten():
                x1, y1, x2, y2 = boxes[i]
                processed_results.append({
                    'class_id': self.classes[class_ids[i]],
                    'confidence': confidences[i],
                    'box': [x1, y1, x2, y2]
                })
        return processed_results

    def infer(self, image,desired_class_id=None):
        h,w,_ = image.shape
        input_image = self.preprocess(image)
        output = self.session.run([self.output_name], {self.input_name: input_image})
        processed_results = self.postprocess(output)
        orig_height, orig_width = image.shape[:2]
        scale_x = orig_width / 640  # 修正:基于模型输入尺寸640计算缩放比例
        scale_y = orig_height / 640
        if desired_class_id is not None:
            boxes = [result['box'] for result in processed_results if result['class_id']==self.classes[desired_class_id]]
            scores = [result['confidence'] for result in processed_results if result['class_id']==self.classes[desired_class_id]]
            class_ids = [result['class_id'] for result in processed_results if result['class_id']==self.classes[desired_class_id]]
        else:
            boxes = [result['box'] for result in processed_results]
            scores = [result['confidence'] for result in processed_results]
            class_ids = [result['class_id'] for result in processed_results]
        return boxes, scores, class_ids, scale_x, scale_y

    def draw_detections(self,image, boxes, scores, class_ids, scale_x, scale_y):
        for box, score, class_id in zip(boxes, scores, class_ids):
            x1, y1, x2, y2 = box

            # Scale the coordinates back to the original image size
            x1 = int(x1 * scale_x)
            y1 = int(y1 * scale_y)
            x2 = int(x2 * scale_x)
            y2 = int(y2 * scale_y)
            
            print(x1, y1, x2, y2)

            # Ensure class_id is used correctly to fetch class name
            class_name = class_id if isinstance(class_id, str) else self.classes[class_id]

            # Draw the bounding box
            cv2.rectangle(image, (x1, y1), (x2, y2), (0, 255, 0), 2)

            # Label the bounding box with class name and confidence score
            label = f'{class_name}: {score:.2f}'
            cv2.putText(image, label, (x1, y1 - 10), cv2.FONT_HERSHEY_SIMPLEX, 0.9, (36, 255, 12), 2)

        return image

if __name__ == "__main__":
    model_path = r"D:\Research licenta\yolo-speedlimits-combined-dataset-nano.onnx"
    image_path = r"D:\AndroidStudioProjects\Licenta\app\src\main\assets\0000000089.png"

    # Load and run inference
    yolo = YoloInference(model_path)
    image = cv2.imread(image_path)
    boxes, scores, class_ids, scale_x, scale_y = yolo.infer(image)

    # Draw and display detections
    image_with_detections = yolo.draw_detections(image, boxes, scores, class_ids, scale_x, scale_y)
    cv2.imwrite("Detectionsl.jpg", image_with_detections)

内容的提问来源于stack exchange,提问作者Gaina-Florin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 05:19:58