You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes中NVIDIA Device Plugin DaemonSet无法运行求助

Kubernetes无法识别NVIDIA GPU问题

问题背景

希望在搭载NVIDIA GPU的服务器/PC上让Kubernetes识别并使用GPU,按照NVIDIA Kubernetes设备插件文档配置后,daemonset无法正常工作。

已完成操作

  1. 安装cri-dockerd
    • 版本验证:cri-dockerd --version
      cri-dockerd 0.2.6 (d8accf7)
      
    • 服务状态验证:systemctl status cri-docker.socket
      ● cri-docker.socket - CRI Docker Socket for the API
           Loaded: loaded (/etc/systemd/system/cri-docker.socket; enabled; vendor preset: enabled)
           Active: active (running) since Mon 2022-12-05 15:00:35 KST; 18h ago
         Triggers: ● cri-docker.service
           Listen: /run/cri-dockerd.sock (Stream)
            Tasks: 0 (limit: 18968)
           Memory: 4.0K
           CGroup: /system.slice/cri-docker.socket
      
      12월 05 15:00:35 hibernation systemd[1]: Starting CRI Docker Socket for the API.
      12월 05 15:00:35 hibernation systemd[1]: Listening on CRI Docker Socket for the API.
      
  2. 安装NVIDIA Docker
    • 功能验证:sudo docker run --rm --gpus all nvidia/cuda:11.3.1-base-ubuntu20.04 nvidia-smi
      +-----------------------------------------------------------------------------+
      | NVIDIA-SMI 470.161.03   Driver Version: 470.161.03   CUDA Version: 11.4     |
      |-------------------------------+----------------------+----------------------+| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC || Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. ||                               |                      |               MIG M. ||===============================+======================+======================||   0  NVIDIA GeForce ...  Off  | 00000000:07:00.0 Off |                  N/A ||  0%   28C    P8    14W / 180W |     64MiB / 12052MiB |      0%      Default ||                               |                      |                  N/A |+-------------------------------+----------------------+----------------------+
                                                                                 
      +-----------------------------------------------------------------------------+
      | Processes:                                                                  ||  GPU   GI   CI        PID   Type   Process name                  GPU Memory ||        ID   ID                                                   Usage      ||=============================================================================|
      +-----------------------------------------------------------------------------+
      
    • 验证/etc/docker/daemon.json配置:
      {
        "default-runtime": "nvidia",
        "runtimes": {
          "nvidia": {
            "path": "nvidia-container-runtime",
            "runtimeArgs": []
          }
        }
      }
      
    • 验证/etc/containerd/config.toml配置:
      # disabled_plugins = ["cri"]        # to annotated
      version = 2
      [plugins]
        [plugins."io.containerd.grpc.v1.cri"]
          [plugins."io.containerd.grpc.v1.cri".containerd]
            default_runtime_name = "nvidia"
      
            [plugins."io.containerd.grpc.v1.cri".containerd.runtimes]
              [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia]
                privileged_without_host_devices = false
                runtime_engine = ""
                runtime_root = ""
                runtime_type = "io.containerd.runc.v2"
                [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia.options]
                  BinaryName = "/usr/bin/nvidia-container-runtime"
      
  3. 安装Kubernetes组件(kubectl=1.22.13-00、kubelet=1.22.13-00、kubeadm=1.22.13-00)
    • 版本验证:kubectl version --client && kubeadm version
      Client Version: version.Info{Major:"1", Minor:"22", GitVersion:"v1.22.13", GitCommit:"a43c0904d0de10f92aa3956c74489c45e6453d6e", GitTreeState:"clean", BuildDate:"2022-08-17T18:28:56Z", GoVersion:"go1.16.15", Compiler:"gc", Platform:"linux/amd64"}
      kubeadm version: &version.Info{Major:"1", Minor:"22", GitVersion:"v1.22.13", GitCommit:"a43c0904d0de10f92aa3956c74489c45e6453d6e", GitTreeState:"clean", BuildDate:"2022-08-17T18:27:51Z", GoVersion:"go1.16.15", Compiler:"gc", Platform:"linux/amd64"}
      
    • kubelet状态验证:systemctl status kubelet
      ● kubelet.service - kubelet: The Kubernetes Node Agent
           Loaded: loaded (/lib/systemd/system/kubelet.service; enabled; vendor preset: enabled)
          Drop-In: /etc/systemd/system/kubelet.service.d
                   └─10-kubeadm.conf
           Active: active (running) since Mon 2022-12-05 15:11:30 KST; 18h ago
             Docs: https://kubernetes.io/docs/home/
      
  4. 初始化Master节点
    • 执行命令:
      sudo kubeadm init \
        --pod-network-cidr=10.244.0.0/16 \
        --apiserver-advertise-address 192.168.219.100\
        --cri-socket /run/cri-dockerd.sock
      
    • 节点状态验证:kubectl get nodes
      NAME             STATUS   ROLES                  AGE     VERSION
      hibernation      Ready    control-plane,master   3m21s   v1.22.13
      

尝试操作

启用Kubernetes GPU支持:

kubectl create -f https://raw.githubusercontent.com/NVIDIA/k8s-device-plugin/v0.13.0/nvidia-device-plugin.yml

问题现象

  • 验证GPU分配情况:kubectl get nodes "-o=custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu"
    预期GPU值为1,实际输出:
    NAME             GPU
    hibernation      <none>
    
  • 执行kubectl get pod -A | grep nvidia无任何输出。
  • 查看daemonset状态:kubectl describe daemonset nvidia-device-plugin-daemonset -n kube-system
    Name:           nvidia-device-plugin-daemonset
    Selector:       name=nvidia-device-plugin-ds
    Node-Selector:  <none>
    Labels:         <none>
    Annotations:    deprecated.daemonset.template.generation: 1
    Desired Number of Nodes Scheduled: 0
    Current Number of Nodes Scheduled: 0
    Number of Nodes Scheduled with Up-to-date Pods: 0
    Number of Nodes Scheduled with Available Pods: 0
    Number of Nodes Misscheduled: 0
    Pods Status:  0 Running / 0 Waiting / 0 Succeeded / 0 Failed
    Pod Template:
      Labels:  name=nvidia-device-plugin-ds
      Containers:
       nvidia-device-plugin-ctr:
        Image:      nvcr.io/nvidia/k8s-device-plugin:v0.13.0
        Port:       <none>
        Host Port:  <none>
        Environment:
          FAIL_ON_INIT_ERROR:  false
        Mounts:
          /var/lib/kubelet/device-plugins from device-plugin (rw)
      Volumes:
       device-plugin:
        Type:               HostPath (bare host directory volume)
        Path:               /var/lib/kubelet/device-plugins
        HostPathType:       
      Priority Class Name:  system-node-critical
    Events:                 <none>
    
    核心异常状态:
    Desired Number of Nodes Scheduled: 0
    Current Number of Nodes Scheduled: 0
    Number of Nodes Scheduled with Up-to-date Pods: 0
    Number of Nodes Scheduled with Available Pods: 0
    

环境信息

Ubuntu 20.04 
GPU: NVIDIA GeForce RTX 3060

本次为格式化桌面后的首次安装,无其他冗余程序。


内容的提问来源于stack exchange,提问作者TaeUk Noh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 00:01:22