You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

YoloNas-l实测推理速度慢于Yolov8-l,求原因分析

YoloNAS-L推理速度慢于YOLOv8-L的原因分析

根据图表数据,YoloNAS-L的推理速度理应远快于YOLOv8-L,但在Google Colab中自行测试却得到相反结果:YoloNAS-L耗时0.4496428966522217,YOLOv8-L耗时0.399261474609375。测试代码如下:

#installing pakcages
!pip install super-gradients==3.1.2
!pip install ultralytics
!pip install pytorch_quantization==2.1.3
!pip install boto3

import super_gradients
from ultralytics import YOLO
import time

!mkdir images
%cd ./images
!wget https://qph.cf2.poecdn.net/main-129957940_34140400692_36_1.png
!wget https://deci-pretrained-models.s3.amazonaws.com/sample_images/beatles-abbeyroad.jpg
!wget https://img.freepik.com/free-photo/view-tiger-animal-wild_23-2150374850.jpg
!wget https://www.fonedog.com/images/photo-compress/image-compressor-image.jpg
!wget https://static.addtoany.com/images/dracaena-cinnabari.jpg
!wget https://qph.cf2.poecdn.net/main-129957940_34139995188_34_1.png
!wget https://qph.cf2.poecdn.net/main-129957940_34139020340_37_1.png

%cd /content/images
import glob
types = ['*.png','*.jpg']
images =[]
for typ in types:
  for image in glob.glob(typ):
    images.append('/content/images/'+image)
    print('/content/images'+image)
nas_ult = NAS('yolo_nas_l.pt')
results = nas_ult.val(data='coco8.yaml')
yolov8 = YOLO("yolov8l.yaml")
yolov8 = YOLO("yolov8l.pt")
nas_inf= 0
nas_res=[]
v8_inf = 0
v8_res=[]
for image in images:
  nas_ult_preds = nas_ult(image,imgsz=640)
for image in images:
  yolov8_predictions = yolov8(image,imgsz=640)
for image in images:
  start = time.time()
  nas_ult_preds = nas_ult(image,imgsz=640)
  end = time.time()
  nas_res.append((nas_ult_preds, end-start))
  nas_inf += end-start
  start = time.time()
  yolov8_predictions = yolov8(image,imgsz=640)
  end = time.time()
  v8_res.append((yolov8_predictions, end-start))
  v8_inf += end-start

print('V8 time:',v8_inf)
print('NAS time:',nas_inf)

核心原因与解决方案

1. 未开启模型推理优化模式

YoloNAS默认可能处于训练模式,而YOLOv8加载后自动进入评估模式。训练模式下会计算梯度、保留中间变量,大幅拖慢推理速度。需要手动开启评估模式和无梯度上下文:

  • 添加nas_ult.model.eval()切换到评估模式
  • 使用torch.no_grad()上下文管理器禁用梯度计算

2. 预热次数不足

虽然你做了两轮预热循环,但仅7张图片的预热不足以让PyTorch完成算子融合、JIT编译等优化。YoloNAS的算子结构更复杂,需要更多预热次数让硬件和框架完成优化。建议增加预热循环到10-20次。

3. 测试样本量过小

仅7张图片的测试结果偶然性极强,单次推理的波动会直接影响总耗时。应该测试更多样本(比如50-100张),并计算单张图片平均耗时,这样结果才具备参考性。

4. 缺少硬件加速配置

YoloNAS需要手动开启TensorRT、ONNX Runtime等硬件加速才能发挥性能优势,而YOLOv8在Colab环境下会自动适配部分加速。可以尝试将YoloNAS导出为ONNX格式后用TensorRT推理,或者使用SuperGradients内置的加速工具。

修正后的测试代码示例

import torch
import super_gradients
from ultralytics import YOLO
import time
import glob

# 安装依赖、下载图片、加载图片步骤略

# 初始化模型并开启优化
nas_ult = super_gradients.training.models.get("yolo_nas_l", pretrained_weights="coco").to('cuda')
nas_ult.model.eval()
yolov8 = YOLO("yolov8l.pt").to('cuda')

# 增加预热次数
warmup_rounds = 15
for _ in range(warmup_rounds):
    nas_ult(images[0], imgsz=640)
    yolov8(images[0], imgsz=640)

# 正式计时,使用无梯度上下文
nas_total = 0
v8_total = 0
test_rounds = 3  # 多轮测试取平均
with torch.no_grad():
    for _ in range(test_rounds):
        for img in images:
            # 测试YoloNAS
            start = time.time()
            nas_ult(img, imgsz=640)
            nas_total += time.time() - start
            
            # 测试YOLOv8
            start = time.time()
            yolov8(img, imgsz=640)
            v8_total += time.time() - start

# 计算平均单张耗时
avg_nas = nas_total / (test_rounds * len(images))
avg_v8 = v8_total / (test_rounds * len(images))

print(f"YOLOv8-L 平均单张耗时: {avg_v8:.6f}s")
print(f"YoloNAS-L 平均单张耗时: {avg_nas:.6f}s")

内容的提问来源于stack exchange,提问作者Reza shahriari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 19:32:32