YoloNas-l实测推理速度慢于Yolov8-l,求原因分析
YoloNAS-L推理速度慢于YOLOv8-L的原因分析
根据图表数据,YoloNAS-L的推理速度理应远快于YOLOv8-L,但在Google Colab中自行测试却得到相反结果:YoloNAS-L耗时0.4496428966522217,YOLOv8-L耗时0.399261474609375。测试代码如下:
#installing pakcages !pip install super-gradients==3.1.2 !pip install ultralytics !pip install pytorch_quantization==2.1.3 !pip install boto3 import super_gradients from ultralytics import YOLO import time !mkdir images %cd ./images !wget https://qph.cf2.poecdn.net/main-129957940_34140400692_36_1.png !wget https://deci-pretrained-models.s3.amazonaws.com/sample_images/beatles-abbeyroad.jpg !wget https://img.freepik.com/free-photo/view-tiger-animal-wild_23-2150374850.jpg !wget https://www.fonedog.com/images/photo-compress/image-compressor-image.jpg !wget https://static.addtoany.com/images/dracaena-cinnabari.jpg !wget https://qph.cf2.poecdn.net/main-129957940_34139995188_34_1.png !wget https://qph.cf2.poecdn.net/main-129957940_34139020340_37_1.png %cd /content/images import glob types = ['*.png','*.jpg'] images =[] for typ in types: for image in glob.glob(typ): images.append('/content/images/'+image) print('/content/images'+image) nas_ult = NAS('yolo_nas_l.pt') results = nas_ult.val(data='coco8.yaml') yolov8 = YOLO("yolov8l.yaml") yolov8 = YOLO("yolov8l.pt") nas_inf= 0 nas_res=[] v8_inf = 0 v8_res=[] for image in images: nas_ult_preds = nas_ult(image,imgsz=640) for image in images: yolov8_predictions = yolov8(image,imgsz=640) for image in images: start = time.time() nas_ult_preds = nas_ult(image,imgsz=640) end = time.time() nas_res.append((nas_ult_preds, end-start)) nas_inf += end-start start = time.time() yolov8_predictions = yolov8(image,imgsz=640) end = time.time() v8_res.append((yolov8_predictions, end-start)) v8_inf += end-start print('V8 time:',v8_inf) print('NAS time:',nas_inf)
核心原因与解决方案
1. 未开启模型推理优化模式
YoloNAS默认可能处于训练模式,而YOLOv8加载后自动进入评估模式。训练模式下会计算梯度、保留中间变量,大幅拖慢推理速度。需要手动开启评估模式和无梯度上下文:
- 添加
nas_ult.model.eval()切换到评估模式 - 使用
torch.no_grad()上下文管理器禁用梯度计算
2. 预热次数不足
虽然你做了两轮预热循环,但仅7张图片的预热不足以让PyTorch完成算子融合、JIT编译等优化。YoloNAS的算子结构更复杂,需要更多预热次数让硬件和框架完成优化。建议增加预热循环到10-20次。
3. 测试样本量过小
仅7张图片的测试结果偶然性极强,单次推理的波动会直接影响总耗时。应该测试更多样本(比如50-100张),并计算单张图片平均耗时,这样结果才具备参考性。
4. 缺少硬件加速配置
YoloNAS需要手动开启TensorRT、ONNX Runtime等硬件加速才能发挥性能优势,而YOLOv8在Colab环境下会自动适配部分加速。可以尝试将YoloNAS导出为ONNX格式后用TensorRT推理,或者使用SuperGradients内置的加速工具。
修正后的测试代码示例
import torch import super_gradients from ultralytics import YOLO import time import glob # 安装依赖、下载图片、加载图片步骤略 # 初始化模型并开启优化 nas_ult = super_gradients.training.models.get("yolo_nas_l", pretrained_weights="coco").to('cuda') nas_ult.model.eval() yolov8 = YOLO("yolov8l.pt").to('cuda') # 增加预热次数 warmup_rounds = 15 for _ in range(warmup_rounds): nas_ult(images[0], imgsz=640) yolov8(images[0], imgsz=640) # 正式计时,使用无梯度上下文 nas_total = 0 v8_total = 0 test_rounds = 3 # 多轮测试取平均 with torch.no_grad(): for _ in range(test_rounds): for img in images: # 测试YoloNAS start = time.time() nas_ult(img, imgsz=640) nas_total += time.time() - start # 测试YOLOv8 start = time.time() yolov8(img, imgsz=640) v8_total += time.time() - start # 计算平均单张耗时 avg_nas = nas_total / (test_rounds * len(images)) avg_v8 = v8_total / (test_rounds * len(images)) print(f"YOLOv8-L 平均单张耗时: {avg_v8:.6f}s") print(f"YoloNAS-L 平均单张耗时: {avg_nas:.6f}s")
内容的提问来源于stack exchange,提问作者Reza shahriari
相关产品推荐
相关产品推荐

