如何让K8s Pod使用节点上指定型号的GPU?
实现指定RTX 3080 Ti型号GPU分配的方案
前提确认
先确认集群中Node Feature Discovery(NFD)已正常运行(GPU Operator默认会部署NFD),执行以下命令查看节点上的GPU型号标签:
kubectl describe node <你的节点名称> | grep "nvidia.com/gpu.product"
正常会输出类似:
nvidia.com/gpu.product=NVIDIA-GeForce-RTX-3080 nvidia.com/gpu.product=NVIDIA-GeForce-RTX-3080-Ti
若节点混插两种GPU,这些标签会同时存在;若不同节点分别部署不同GPU,单个节点只会有对应型号的标签。
方案1:节点标签+Pod nodeSelector指定(适合单节点单GPU型号场景)
如果RTX 3080和3080 Ti分别在不同节点上,直接给对应节点打标签后,在Pod中指定选择该节点:
- 给搭载RTX 3080 Ti的节点打自定义标签(也可直接用NFD自动生成的标签):
kubectl label node <目标节点名> gpu-model=rtx3080ti
- 编写Pod配置,通过
nodeSelector指定节点,同时请求GPU资源:
apiVersion: v1 kind: Pod metadata: name: gpu-test-pod spec: containers: - name: gpu-container image: nvidia/cuda:11.8.0-runtime-ubuntu22.04 command: ["sleep", "infinity"] resources: limits: nvidia.com/gpu: 1 # 请求1块GPU nodeSelector: gpu-model: rtx3080ti # 选择带该标签的节点 # 也可直接用NFD原生标签:nvidia.com/gpu.product=NVIDIA-GeForce-RTX-3080-Ti
方案2:Device Plugin配置按型号请求资源(适合单节点混插多GPU型号场景)
如果同一节点上同时插了RTX 3080和3080 Ti,需修改GPU Operator的ClusterPolicy,配置Device Plugin支持按型号区分资源:
- 编辑ClusterPolicy资源:
kubectl edit clusterpolicy nvidia-gpu-operator
- 在
spec.devicePlugin下添加配置,按GPU型号生成专属资源名:
spec: devicePlugin: config: name: nvidia-device-plugin-config data: | version: v1 flags: - --mig-strategy=none resources: - name: nvidia.com/rtx3080ti selectors: - product: "NVIDIA GeForce RTX 3080 Ti" - name: nvidia.com/rtx3080 selectors: - product: "NVIDIA GeForce RTX 3080"
注意product的值要和NFD输出的GPU型号名称完全一致。
保存配置后,GPU Operator会自动重启Device Plugin组件。
编写Pod配置,直接请求RTX 3080 Ti专属资源:
apiVersion: v1 kind: Pod metadata: name: gpu-test-ti-pod spec: containers: - name: gpu-container image: nvidia/cuda:11.8.0-runtime-ubuntu22.04 command: ["sleep", "infinity"] resources: limits: nvidia.com/rtx3080ti: 1 # 直接请求RTX 3080 Ti型号GPU
验证方法
Pod启动后,进入容器执行以下命令确认GPU型号:
kubectl exec -it gpu-test-pod -- nvidia-smi
输出中显示GPU型号为RTX 3080 Ti,即配置生效。
内容的提问来源于stack exchange,提问作者rxopy
相关产品推荐
相关产品推荐

