默认性能更优时,为何还要使用Nvidia MPS、Time Slicing或MIG?
关于Nvidia GPU共享策略(MIG/Time Slicing/MPS)的性能疑惑与测试分析
测试背景
为明确Nvidia三种GPU共享策略(MIG、Time Slicing、MPS)的实际影响,在K8s环境下部署7个相同应用副本,采用80GB显存的A100 GPU,分别开展YOLOv8推理和大矩阵乘法两类基准测试,对比三种策略与默认设置的表现。
测试代码
YOLOv8推理测试代码
import os # Set YOLOv8 to quiet mode os.environ['YOLO_VERBOSE'] = 'False' from prometheus_client import start_http_server, Histogram from ultralytics import YOLO import torch start_http_server(8000) device = torch.device("cuda") model = YOLO("yolov8n.pt").to(device=device) h = Histogram('gpu_stress_inference_yolov8_milliseconds_duration', 'Description of histogram', buckets=(1, 5, 10, 15, 20, 25, 30, 35, 40, 50, 75, 100, 150, 200, 500, 1000, 5000)) def run_model(): results = model("https://ultralytics.com/images/bus.jpg") # print(model.device.type) h.observe(results[0].speed['inference']) while True: run_model()
大矩阵乘法测试代码
import torch import time from prometheus_client import start_http_server, Histogram # Check if CUDA is available and Tensor Cores are supported if not torch.cuda.is_available(): raise SystemError("CUDA is not available on this system") device = torch.device("cuda") torch.cuda.set_sync_debug_mode(debug_mode="warn") torch.set_default_device(device) # ensure we actually use the GPU and don't do the calculations on the CPU h = Histogram('gpu_stress_mat_mul_seconds_duration', 'Description of histogram', buckets=(0.001, 0.005, 0.01, 0.1, 0.25, 0.5, 1.0, 2.0, 3.0, 4.0, 5.0, 10.0, 20.0, 50.0, 100.0, 200.0, 500.0, 1000.0)) def mat_mul(m1, m2): return torch.matmul(m1, m2) # Function to perform matrix multiplication using Tensor Cores def stress(matrix_size=16384): # Create random matrices on the GPU m1 = torch.randn(matrix_size, matrix_size, dtype=torch.float16) m2 = torch.randn(matrix_size, matrix_size, dtype=torch.float16) # Perform matrix multiplications indefinitely while True: start = time.time() output = torch.matmul(m1, m2) print(output.any()) end = time.time() h.observe(end - start) if __name__ == "__main__": start_http_server(8000) stress()
测试结果
- YOLOv8推理场景:默认设置的推理速度最优,且默认设置已支持多应用共享GPU;
- 大矩阵乘法场景:四种策略(MIG/Time Slicing/MPS/默认)的性能无明显差异。
核心疑惑与解答
已知MIG具备内存隔离特性,但如果优先考虑性能,除吞吐量优势外,是否还有必要使用这些共享策略?这与Nvidia文档中提及的策略优势存在矛盾。
关键分析点
默认共享的本质:CUDA默认的多进程共享GPU,本质是基于时间切片的硬件级调度,对于计算负载均匀、无严重资源争抢的场景(如7个YOLOv8副本、大矩阵乘法),这种调度的开销极低,甚至优于显式配置的MPS/Time Slicing——因为显式策略可能引入额外的调度层或资源划分损耗。
共享策略的适用场景:
- MIG:核心价值是硬隔离,适合多租户场景(比如不同团队/用户共享GPU),避免某一应用耗尽显存影响其他应用;当场景不需要隔离,仅追求性能时,MIG的资源划分(比如拆分GPU为多个实例)反而会降低单应用能使用的算力,导致性能下降。
- MPS:针对计算密集型、多进程小任务优化,能减少CUDA上下文切换开销,适合大量小批量推理或计算的场景;但YOLOv8是单循环推理,矩阵乘法是大任务,MPS的优势无法体现,甚至可能因为调度逻辑增加额外开销。
- Time Slicing(K8s设备插件层面):主要是为了实现GPU的细粒度资源配额(比如按百分比分配GPU),适合需要精确控制每个Pod GPU使用率的场景;如果负载不需要配额限制,默认的硬件时间切片效率更高。
文档与实际差异的原因:Nvidia文档强调的是策略的特定场景优势,而非所有场景都优于默认设置。默认共享是最轻量化的方式,在负载匹配时性能最优;而其他策略是为了解决特定问题(隔离、配额、小任务调度)而设计,牺牲部分性能来换取功能特性。
结论
如果场景无需多租户隔离、无需精确GPU资源配额、负载为大任务/均匀负载,优先使用默认共享策略即可,性能表现最优;只有当需要对应策略的特性(如MIG的隔离、MPS的小任务吞吐量、Time Slicing的配额)时,才需要权衡性能损耗启用这些策略。
内容的提问来源于stack exchange,提问作者Ronan Quigley
相关产品推荐
相关产品推荐

