You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并两个YOLOv5模型以消除NumPy与Tensor互转的性能开销

双YOLOv5级联推理优化方案

核心问题根因

你当前使用的YOLOv5官方results.crop()接口默认会把GPU上的检测结果转成CPU侧的NumPy数组返回,这就导致了多次CPU-GPU数据拷贝和格式转换的额外开销。

优化实现(全程GPU侧运算,无NumPy转换)

前置配置

  • 加载模型时直接指定运行设备为CUDA,避免模型权重自动迁移的开销
  • 输入帧提前转换为GPU侧张量,所有裁剪、resize操作都用PyTorch原生接口在GPU完成

优化后代码

import torch
import cv2
import numpy as np

# 加载模型时指定GPU设备
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = torch.hub.load('.', 'custom', path=img_cls_path, source='local', force_reload=True, device=device)
model_ocr = torch.hub.load('.', 'custom', path=ocr_path, source='local', force_reload=True, device=device)
# 固定第二个模型的输入尺寸,提前拿到方便后续resize
ocr_input_size = model_ocr.img_size[0]

cap = cv2.VideoCapture(some_video_path)

while cap.isOpened():
    ret, frame = cap.read()
    if not ret:
        break
    
    # 先把BGR的numpy帧转成GPU侧RGB CHW张量,归一化到0-1,符合YOLOv5输入要求
    frame_tensor = torch.from_numpy(frame).to(device)
    frame_tensor = frame_tensor.permute(2, 0, 1).flip(0) # HWC->CHW, BGR->RGB
    frame_tensor = frame_tensor / 255.0
    # 加batch维度
    frame_tensor = frame_tensor.unsqueeze(0)
    
    # 第一个模型推理,直接拿GPU上的预测结果
    results = model(frame_tensor)
    # pred格式:[x1, y1, x2, y2, confidence, class_id],全部在GPU上
    pred = results.pred[0]
    
    # 过滤出number类的检测框
    number_cls_id = [k for k, v in model.names.items() if v == 'number'][0]
    number_boxes = pred[pred[:, 5] == number_cls_id]
    
    for box in number_boxes:
        x1, y1, x2, y2 = map(int, box[:4])
        # 边界校验,避免越界
        x1, y1 = max(0, x1), max(0, y1)
        x2, y2 = min(frame_tensor.shape[3], x2), min(frame_tensor.shape[2], y2)
        # GPU侧直接裁剪张量
        crop_tensor = frame_tensor[:, :, y1:y2, x1:x2]
        # GPU侧做resize适配第二个模型的输入尺寸,用bilinear插值和原图对齐
        crop_tensor = torch.nn.functional.interpolate(crop_tensor, size=(ocr_input_size, ocr_input_size), mode='bilinear', align_corners=False)
        
        # 直接传GPU张量给第二个模型推理,无格式转换开销
        ocr_results = model_ocr(crop_tensor)
        ocr_pred = ocr_results.pred[0]
        # 后续处理ocr_pred即可,全程在GPU

端到端合并模型方案

如果需要将两个模型合并成单个可部署的模型,可按以下步骤操作:

  • 定义一个继承自torch.nn.Module的类,将两个YOLOv5模型作为子模块
  • 在forward方法中实现上述级联推理、裁剪、resize的逻辑
  • 合并后的模型可以直接导出为ONNX、TensorRT等格式,适合生产环境部署
class CascadedYOLOv5(torch.nn.Module):
    def __init__(self, det_model, ocr_model, ocr_input_size=640):
        super().__init__()
        self.det_model = det_model
        self.ocr_model = ocr_model
        self.ocr_input_size = ocr_input_size
        self.number_cls_id = [k for k, v in det_model.names.items() if v == 'number'][0]
    
    def forward(self, x):
        # x为输入的batch张量,GPU侧
        det_pred = self.det_model(x).pred[0]
        number_boxes = det_pred[det_pred[:, 5] == self.number_cls_id]
        ocr_outputs = []
        for box in number_boxes:
            x1, y1, x2, y2 = map(int, box[:4])
            x1, y1 = max(0, x1), max(0, y1)
            x2, y2 = min(x.shape[3], x2), min(x.shape[2], y2)
            crop = x[:, :, y1:y2, x1:x2]
            crop = torch.nn.functional.interpolate(crop, size=(self.ocr_input_size, self.ocr_input_size), mode='bilinear', align_corners=False)
            ocr_out = self.ocr_model(crop)
            ocr_outputs.append(ocr_out)
        return number_boxes, ocr_outputs

# 实例化合并模型
cascaded_model = CascadedYOLOv5(model, model_ocr)
cascaded_model = cascaded_model.to(device)

注意事项

  • 裁剪时要做边界校验,避免出现索引越界的报错
  • 若第二个模型对输入有均值方差归一化的要求,直接在GPU侧对裁剪后的张量做运算即可,无需转NumPy
  • 批量处理场景下可对裁剪后的多个区域做padding后拼成batch输入第二个模型,进一步提升推理效率

内容的提问来源于stack exchange,提问作者Nurislom Rakhmatullaev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 00:30:01