执行helm upgrade后Pod处于Init:0/2状态及沙箱创建错误求助
修改config.yaml后执行helm upgrade命令,发现两个Pod(continuous-image-puller-2qh7z、hook-image-puller-sz8qf)卡在Init:0/2状态,反复触发FailedCreatePodSandBox错误,核心报错为Multus组件获取Pod时出现Unauthorized。
kubectl get pods 输出
NAME READY STATUS RESTARTS AGE continuous-image-puller-2hhv4 1/1 Running 0 30h continuous-image-puller-2qh7z 0/1 Init:0/2 0 30h hook-image-awaiter-gg7n8 1/1 Running 0 17s hook-image-puller-sz8qf 0/1 Init:0/2 0 18s hook-image-puller-xnc65 1/1 Running 0 18s hub-567bb6bc68-g5lcs 1/1 Running 0 30h ngshare-586b877ff5-wg8sq 0/1 Terminating 0 86d ngshare-85c7799db8-gw6hq 1/1 Running 0 35d proxy-65d966d5b6-cz2rg 1/1 Running 0 30h
kubectl describe pods continuous-image-puller-2qh7z 输出
Name: continuous-image-puller-2qh7z Namespace: staging-jhub Priority: 0 Node: star11/172.27.188.111 Start Time: Mon, 05 Jun 2023 17:03:51 -0700 Labels: app=jupyterhub component=continuous-image-puller controller-revision-hash=84bfb4599d pod-template-generation=4 release=staging-jhub Annotations: <none> Status: Pending IP: IPs: <none> Controlled By: DaemonSet/continuous-image-puller Init Containers: image-pull-metadata-block: Container ID: Image: jupyterhub/k8s-network-tools:2.0.1-0.dev.git.5866.h7de20b77 Image ID: Port: <none> Host Port: <none> Command: /bin/sh -c echo "Pulling complete" State: Waiting Reason: PodInitializing Ready: False Restart Count: 0 Environment: <none> Mounts: <none> image-pull-singleuser: Container ID: Image: libretextsregistry.ddns.net/binder-dev-libretexts-2ddefault-2denv-1cb626:4895042a71a713052ffbf05a0fe907098bf368ab Image ID: Port: <none> Host Port: <none> Command: /bin/sh -c echo "Pulling complete" State: Waiting Reason: PodInitializing Ready: False Restart Count: 0 Environment: <none> Mounts: <none> Containers: pause: Container ID: Image: registry.k8s.io/pause:3.9 Image ID: Port: <none> Host Port: <none> State: Waiting Reason: PodInitializing Ready: False Restart Count: 0 Environment: <none> Mounts: <none> Conditions: Type Status Initialized False Ready False ContainersReady False PodScheduled True Volumes: <none> QoS Class: BestEffort Node-Selectors: <none> Tolerations: hub.jupyter.org/dedicated=user:NoSchedule hub.jupyter.org_dedicated=user:NoSchedule node.kubernetes.io/disk-pressure:NoSchedule op=Exists node.kubernetes.io/memory-pressure:NoSchedule op=Exists node.kubernetes.io/not-ready:NoExecute op=Exists node.kubernetes.io/pid-pressure:NoSchedule op=Exists node.kubernetes.io/unreachable:NoExecute op=Exists node.kubernetes.io/unschedulable:NoSchedule op=Exists Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning FailedCreatePodSandBox 2m10s (x8443 over 30h) kubelet (combined from similar events): Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "563235750ab2900720f155529383e7e8c6e36af111d850ddbefa76b5ad60dfa3": Multus: [staging-jhub/continuous-image-puller-2qh7z]: error getting pod: Unauthorized
kubectl describe pods hook-image-puller-sz8qf 输出
Name: hook-image-puller-sz8qf Namespace: staging-jhub Priority: 0 Node: star11/172.27.188.111 Start Time: Tue, 06 Jun 2023 23:43:02 -0700 Labels: app=jupyterhub component=hook-image-puller controller-revision-hash=b8b77fc77 pod-template-generation=1 release=staging-jhub Annotations: <none> Status: Pending IP: IPs: <none> Controlled By: DaemonSet/hook-image-puller Init Containers: image-pull-metadata-block: Container ID: Image: jupyterhub/k8s-network-tools:2.0.0 Image ID: Port: <none> Host Port: <none> Command: /bin/sh -c echo "Pulling complete" State: Waiting Reason: PodInitializing Ready: False Restart Count: 0 Environment: <none> Mounts: <none> image-pull-singleuser: Container ID: Image: libretextsregistry.ddns.net/binder-dev-libretexts-2ddefault-2denv-1cb626:230d85d708cb6f6d1b6ba36bd2919a431eede445 Image ID: Port: <none> Host Port: <none> Command: /bin/sh -c echo "Pulling complete" State: Waiting Reason: PodInitializing Ready: False Restart Count: 0 Environment: <none> Mounts: <none> Containers: pause: Container ID: Image: k8s.gcr.io/pause:3.8 Image ID: Port: <none> Host Port: <none> State: Waiting Reason: PodInitializing Ready: False Restart Count: 0 Environment: <none> Mounts: <none> Conditions: Type Status Initialized False Ready False ContainersReady False PodScheduled True Volumes: <none> QoS Class: BestEffort Node-Selectors: <none> Tolerations: hub.jupyter.org/dedicated=user:NoSchedule hub.jupyter.org_dedicated=user:NoSchedule node.kubernetes.io/disk-pressure:NoSchedule op=Exists node.kubernetes.io/memory-pressure:NoSchedule op=Exists node.kubernetes.io/not-ready:NoExecute op=Exists node.kubernetes.io/pid-pressure:NoSchedule op=Exists node.kubernetes.io/unreachable:NoExecute op=Exists node.kubernetes.io/unschedulable:NoSchedule op=Exists Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal Scheduled 5m2s default-scheduler Successfully assigned staging-jhub/hook-image-puller-sz8qf to star11 Warning FailedCreatePodSandBox 5m2s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "480a19ef1acceddd4ac6397ce662b089f1ff7234cdc20cc3a24f3667125b7779": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 4m48s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "d2e751acddb353b5d819eb02f58169684f2300f3a25dc86d788a5338ff915c21": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 4m37s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "d3a7e3b0d2a33d97fe80bf1bd24564f10d8d6cf522ec5675de6b49a9c295be7b": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 4m22s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "7b208eea6c900d2af63e003e62094bf2b918d1f8d317d98ff9eb9dd763d119ad": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 4m10s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "0d4088512fa8fe86bf73b921d3d4b35ead6b1389a1b2c86ba9fd0c6bf261d62d": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 3m59s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "2de035d394b5d5d51215262c95ff584b4c133db42f30ec9c7fb43b64a8ab6ac3": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 3m44s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "606274f002f03a444b7549ccb85e61a64017aebf649605994bed0cf15f84e518": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 3m31s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "d60db7015325f37f1fcfa6438845515786e49d9d599f6640dba5b8a1d3bc407d": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 3m18s kubelet Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "cd70873b3b424c659093b9f948fa0717b9680d5935bfbb555e2ef194243b0830": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized Warning FailedCreatePodSandBox 11s (x14 over 3m4s) kubelet (combined from similar events): Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "d404a995dc32e2780533c0ce0afc1f8e5681f8ddfaca289c082b088da9e72728": Multus: [staging-jhub/hook-image-puller-sz8qf]: error getting pod: Unauthorized
核心问题定位
两个异常Pod都调度到star11节点,且均出现Multus网络插件访问Pod资源时未授权的错误,说明该节点上的Multus组件权限配置异常,或kubelet与Multus的通信凭证失效。
具体排查步骤
检查Multus服务账户权限:执行以下命令,确认Multus使用的ServiceAccount绑定的ClusterRole包含
pods资源的get权限:kubectl describe sa multus -n kube-system kubectl describe clusterrolebinding multus验证节点Multus配置:登录
star11节点,检查/etc/cni/net.d/下的Multus配置文件,确认配置中引用的kubeconfig凭证有效,或是否指向正确的ServiceAccount令牌路径。查看kubelet日志:检查
star11节点的kubelet日志,排查是否有凭证过期、权限变更相关记录:journalctl -u kubelet -f重启Multus组件:若为临时权限失效,可尝试重启节点上的Multus容器,或通过DaemonSet重启整个Multus部署:
# 节点上直接重启容器 docker restart $(docker ps | grep multus | awk '{print $1}') # 通过DaemonSet重启 kubectl rollout restart daemonset multus -n kube-system核对Helm升级变更:对比升级前后的
config.yaml,确认是否修改了网络策略、Multus配置参数,或节点容忍度/标签,导致Pod调度到权限异常的节点。重新调度异常Pod:删除异常Pod,让DaemonSet重新创建:
kubectl delete pod continuous-image-puller-2qh7z hook-image-puller-sz8qf -n staging-jhub
常见解决方法
- 若Multus的ServiceAccount缺少权限,创建或更新ClusterRole,添加
pods资源的get权限,再重新绑定到ServiceAccount。 - 若节点上的kubeconfig凭证过期,替换为当前有效的ServiceAccount令牌(路径通常为
/var/run/secrets/kubernetes.io/serviceaccount/token)。 - 若Helm升级误改网络配置,回滚到之前的稳定版本:
helm rollback staging-jhub <目标版本号>
内容的提问来源于stack exchange,提问作者user21894767

