能否显示Kubernetes镜像拉取进度?Pod卡Pending状态求助
Kubernetes Pod卡在Pending(ContainerCreating)问题排查与解决
问题场景
执行Deployment滚动更新时,Pod卡在Pending状态超过3分钟(正常镜像拉取耗时<1分钟)。kubectl describe pod输出如下:
➜ retire git:(gateway) kubectl describe pod dolphin-post-service-6d8b4fd8b4-nl4kd -n reddwarf-pro Name: dolphin-post-service-6d8b4fd8b4-nl4kd Namespace: reddwarf-pro Priority: 0 Node: k8smasterone/172.29.217.209 Start Time: Wed, 10 Aug 2022 00:33:23 +0800 Labels: app=dolphin-post-service pod-template-hash=6d8b4fd8b4 Annotations: cni.projectcalico.org/containerID: fb54cff6a7fa39419f3249f89e8fd9e41768182bb3917aef0183452ff331789c cni.projectcalico.org/podIP: 10.97.196.243/32 cni.projectcalico.org/podIPs: 10.97.196.243/32 kubectl.kubernetes.io/restartedAt: 2022-07-22T14:49:57Z Status: Pending IP: IPs: <none> Controlled By: ReplicaSet/dolphin-post-service-6d8b4fd8b4 Containers: dolphin-post-service: Container ID: Image: registry.cn-hongkong.aliyuncs.com/reddwarf-pro/dolphin-post:cc5cd79c4aa5bf9f0a2aba0c9a775be7e6d3d469 Image ID: Port: 80/TCP Host Port: 0/TCP State: Waiting Reason: ContainerCreating Ready: False Restart Count: 0 Limits: cpu: 1500M memory: 600Mi Requests: cpu: 200m memory: 256Mi Liveness: http-get http://:11014/actuator/health/liveness delay=190s timeout=30s period=30s #success=1 #failure=3 Readiness: http-get http://:11014/actuator/health delay=160s timeout=30s period=30s #success=1 #failure=3 Environment: APOLLO_META: <set to the key 'apollo.meta' of config map 'pro-apollo-config'> Optional: false ENV: <set to the key 'env' of config map 'pro-apollo-config'> Optional: false Mounts: /root/data from dolphin-post-service-persistent-storage (rw) /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-l6m2z (ro) Conditions: Type Status Initialized True Ready False ContainersReady False PodScheduled True Volumes: dolphin-post-service-persistent-storage: Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace) ClaimName: dolphin-post-service-pv-claim ReadOnly: false kube-api-access-l6m2z: Type: Projected (a volume that contains injected data from multiple sources) TokenExpirationSeconds: 3607 ConfigMapName: kube-root-ca.crt ConfigMapOptional: <nil> DownwardAPI: true QoS Class: Burstable Node-Selectors: <none> Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s node.kubernetes.io/unreachable:NoExecute op=Exists for 300s Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal Scheduled 3m53s default-scheduler Successfully assigned reddwarf-pro/dolphin-post-service-6d8b4fd8b4-nl4kd to k8smasterone Normal Pulling 3m52s kubelet Pulling image "registry.cn-hongkong.aliyuncs.com/reddwarf-pro/dolphin-post:cc5cd79c4aa5bf9f0a2aba0c9a775be7e6d3d469"
尝试查看容器日志时报错:
Failed to load logs: container "dolphin-post-service" in pod "dolphin-post-service-6d8b4fd8b4-nl4kd" is waiting to start: ContainerCreating Reason: BadRequest (400)
可能原因
从Pod事件和状态来看,核心问题是镜像拉取停滞,常见原因包括:
- 节点与阿里云香港镜像仓库的网络连通性差:比如节点所在网络到香港地域的带宽不足、防火墙拦截了镜像仓库的443端口
- 镜像拉取权限缺失:节点或Deployment未配置阿里云镜像仓库的拉取凭证,无法访问私有镜像
- 镜像不存在或标签错误:指定的镜像标签
cc5cd79c4aa5bf9f0a2aba0c9a775be7e6d3d469在仓库中不存在,或镜像名称拼写错误 - 节点磁盘空间不足:节点本地存储已满,无法存储拉取的镜像层
- 镜像过大或网络波动:导致拉取超时,超出预期耗时
解决方法
1. 手动验证节点镜像拉取能力
登录Pod调度到的节点k8smasterone,执行手动拉取命令,直接查看错误信息:
docker pull registry.cn-hongkong.aliyuncs.com/reddwarf-pro/dolphin-post:cc5cd79c4aa5bf9f0a2aba0c9a775be7e6d3d469
根据报错提示针对性处理:
- 权限错误:配置阿里云镜像仓库的凭证
- 网络超时:检查节点网络,或切换到国内地域的镜像仓库
- 镜像不存在:确认镜像标签正确性
2. 检查节点网络连通性
在节点上测试与镜像仓库的连通性:
# 测试HTTPS端口是否可达 telnet registry.cn-hongkong.aliyuncs.com 443 # 或用curl测试 curl -I https://registry.cn-hongkong.aliyuncs.com
如果无法连通,排查节点网络策略、防火墙规则,或更换镜像仓库地域。
3. 配置镜像拉取凭证
如果是私有镜像,确保Deployment配置了imagePullSecrets:
- 创建阿里云镜像仓库的Secret:
kubectl create secret docker-registry aliyun-reg-secret \ --docker-server=registry.cn-hongkong.aliyuncs.com \ --docker-username=<你的阿里云账号> \ --docker-password=<你的镜像仓库密码> \ --namespace=reddwarf-pro
- 在Deployment的
spec.template.spec中添加:
imagePullSecrets: - name: aliyun-reg-secret
4. 检查节点资源
查看节点磁盘空间和内存:
# 磁盘空间 df -h # 内存使用 free -m
如果磁盘不足,清理节点上的无用镜像或日志,释放空间后重新触发Pod更新。
5. 查看kubelet详细日志
在节点上查看kubelet的实时日志,获取镜像拉取的详细错误:
journalctl -u kubelet -f
日志中会包含镜像拉取的具体失败原因,比如超时、权限拒绝等。
镜像拉取进度查看
Kubernetes原生的kubectl命令无法直接查看镜像拉取的实时进度,因为kubelet不会向API Server上报拉取进度。可以通过以下两种方式查看:
- 手动在节点拉取:执行
docker pull <镜像地址>,终端会显示每层镜像的下载进度 - 查看kubelet日志:kubelet日志中会输出镜像拉取的步骤,比如"Downloaded layer xxx"等信息,间接判断拉取进度
内容的提问来源于stack exchange,提问作者Dolphin
相关产品推荐
相关产品推荐

