Kubernetes中NVIDIA Device Plugin DaemonSet无法运行求助
Kubernetes无法识别NVIDIA GPU问题
问题背景
希望在搭载NVIDIA GPU的服务器/PC上让Kubernetes识别并使用GPU,按照NVIDIA Kubernetes设备插件文档配置后,daemonset无法正常工作。
已完成操作
- 安装cri-dockerd
- 版本验证:
cri-dockerd --versioncri-dockerd 0.2.6 (d8accf7) - 服务状态验证:
systemctl status cri-docker.socket● cri-docker.socket - CRI Docker Socket for the API Loaded: loaded (/etc/systemd/system/cri-docker.socket; enabled; vendor preset: enabled) Active: active (running) since Mon 2022-12-05 15:00:35 KST; 18h ago Triggers: ● cri-docker.service Listen: /run/cri-dockerd.sock (Stream) Tasks: 0 (limit: 18968) Memory: 4.0K CGroup: /system.slice/cri-docker.socket 12월 05 15:00:35 hibernation systemd[1]: Starting CRI Docker Socket for the API. 12월 05 15:00:35 hibernation systemd[1]: Listening on CRI Docker Socket for the API.
- 版本验证:
- 安装NVIDIA Docker
- 功能验证:
sudo docker run --rm --gpus all nvidia/cuda:11.3.1-base-ubuntu20.04 nvidia-smi+-----------------------------------------------------------------------------+ | NVIDIA-SMI 470.161.03 Driver Version: 470.161.03 CUDA Version: 11.4 | |-------------------------------+----------------------+----------------------+| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC || Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. || | | MIG M. ||===============================+======================+======================|| 0 NVIDIA GeForce ... Off | 00000000:07:00.0 Off | N/A || 0% 28C P8 14W / 180W | 64MiB / 12052MiB | 0% Default || | | N/A |+-------------------------------+----------------------+----------------------+ +-----------------------------------------------------------------------------+ | Processes: || GPU GI CI PID Type Process name GPU Memory || ID ID Usage ||=============================================================================| +-----------------------------------------------------------------------------+ - 验证
/etc/docker/daemon.json配置:{ "default-runtime": "nvidia", "runtimes": { "nvidia": { "path": "nvidia-container-runtime", "runtimeArgs": [] } } } - 验证
/etc/containerd/config.toml配置:# disabled_plugins = ["cri"] # to annotated version = 2 [plugins] [plugins."io.containerd.grpc.v1.cri"] [plugins."io.containerd.grpc.v1.cri".containerd] default_runtime_name = "nvidia" [plugins."io.containerd.grpc.v1.cri".containerd.runtimes] [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia] privileged_without_host_devices = false runtime_engine = "" runtime_root = "" runtime_type = "io.containerd.runc.v2" [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia.options] BinaryName = "/usr/bin/nvidia-container-runtime"
- 功能验证:
- 安装Kubernetes组件(kubectl=1.22.13-00、kubelet=1.22.13-00、kubeadm=1.22.13-00)
- 版本验证:
kubectl version --client && kubeadm versionClient Version: version.Info{Major:"1", Minor:"22", GitVersion:"v1.22.13", GitCommit:"a43c0904d0de10f92aa3956c74489c45e6453d6e", GitTreeState:"clean", BuildDate:"2022-08-17T18:28:56Z", GoVersion:"go1.16.15", Compiler:"gc", Platform:"linux/amd64"} kubeadm version: &version.Info{Major:"1", Minor:"22", GitVersion:"v1.22.13", GitCommit:"a43c0904d0de10f92aa3956c74489c45e6453d6e", GitTreeState:"clean", BuildDate:"2022-08-17T18:27:51Z", GoVersion:"go1.16.15", Compiler:"gc", Platform:"linux/amd64"} - kubelet状态验证:
systemctl status kubelet● kubelet.service - kubelet: The Kubernetes Node Agent Loaded: loaded (/lib/systemd/system/kubelet.service; enabled; vendor preset: enabled) Drop-In: /etc/systemd/system/kubelet.service.d └─10-kubeadm.conf Active: active (running) since Mon 2022-12-05 15:11:30 KST; 18h ago Docs: https://kubernetes.io/docs/home/
- 版本验证:
- 初始化Master节点
- 执行命令:
sudo kubeadm init \ --pod-network-cidr=10.244.0.0/16 \ --apiserver-advertise-address 192.168.219.100\ --cri-socket /run/cri-dockerd.sock - 节点状态验证:
kubectl get nodesNAME STATUS ROLES AGE VERSION hibernation Ready control-plane,master 3m21s v1.22.13
- 执行命令:
尝试操作
启用Kubernetes GPU支持:
kubectl create -f https://raw.githubusercontent.com/NVIDIA/k8s-device-plugin/v0.13.0/nvidia-device-plugin.yml
问题现象
- 验证GPU分配情况:
kubectl get nodes "-o=custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu"
预期GPU值为1,实际输出:NAME GPU hibernation <none> - 执行
kubectl get pod -A | grep nvidia无任何输出。 - 查看daemonset状态:
kubectl describe daemonset nvidia-device-plugin-daemonset -n kube-system
核心异常状态:Name: nvidia-device-plugin-daemonset Selector: name=nvidia-device-plugin-ds Node-Selector: <none> Labels: <none> Annotations: deprecated.daemonset.template.generation: 1 Desired Number of Nodes Scheduled: 0 Current Number of Nodes Scheduled: 0 Number of Nodes Scheduled with Up-to-date Pods: 0 Number of Nodes Scheduled with Available Pods: 0 Number of Nodes Misscheduled: 0 Pods Status: 0 Running / 0 Waiting / 0 Succeeded / 0 Failed Pod Template: Labels: name=nvidia-device-plugin-ds Containers: nvidia-device-plugin-ctr: Image: nvcr.io/nvidia/k8s-device-plugin:v0.13.0 Port: <none> Host Port: <none> Environment: FAIL_ON_INIT_ERROR: false Mounts: /var/lib/kubelet/device-plugins from device-plugin (rw) Volumes: device-plugin: Type: HostPath (bare host directory volume) Path: /var/lib/kubelet/device-plugins HostPathType: Priority Class Name: system-node-critical Events: <none>Desired Number of Nodes Scheduled: 0 Current Number of Nodes Scheduled: 0 Number of Nodes Scheduled with Up-to-date Pods: 0 Number of Nodes Scheduled with Available Pods: 0
环境信息
Ubuntu 20.04 GPU: NVIDIA GeForce RTX 3060
本次为格式化桌面后的首次安装,无其他冗余程序。
内容的提问来源于stack exchange,提问作者TaeUk Noh
相关产品推荐
相关产品推荐

