Kube-dns始终处于Pending状态求助(kubeadm部署K8s集群)
Hey there, let's walk through the most common reasons why your kube-dns Pod is stuck in Pending after adding a new node to your kubeadm-provisioned K8s cluster on virt-manager VMs. I'll break down actionable checks and fixes:
1. Check for Insufficient Node Resources
The most frequent culprit is that none of your nodes have enough CPU/memory to run the kube-dns Pod. To confirm:
- Run this command to inspect the Pod's details and events:
kubectl describe pod kube-dns-86f4d74b45-br6ck -n kube-system - Look in the Events section for messages like
Insufficient cpuorInsufficient memory.
Fix:
- Allocate more CPU/memory to your virt-manager VMs (edit the VM settings to increase resources).
- Check if other Pods are hogging resources with
kubectl top pods --all-namespacesand scale them back if possible.
2. Verify Node Readiness
If your new node isn't marked as Ready, the scheduler can't assign kube-dns to it.
- Check node statuses:
kubectl get nodes - If the new node shows
NotReady, dig deeper:- Check the kubelet service status on the new node:
systemctl status kubelet - View kubelet logs for errors:
journalctl -u kubelet -f - Ensure firewall rules on the node allow K8s traffic (e.g., 6443, 10250, 10256 ports are open) and SELinux is disabled or configured properly.
- Check the kubelet service status on the new node:
3. Confirm CNI Network Plugin is Installed and Working
Kubernetes requires a CNI (Container Network Interface) plugin to enable Pod-to-Pod communication. kubeadm doesn't install one by default, so missing CNI will leave Pods stuck in Pending.
- Check if CNI-related Pods are running:
kubectl get pods -n kube-system | grep -i cni - If you don't see any running CNI Pods (like flannel, calico, or weave), install a compatible plugin. For example, to set up Flannel, apply its manifest file (ensure it matches your K8s version):
kubectl apply -f <path-to-local-flannel-manifest> - Also, verify the CNI config exists on all nodes:
This directory should have als /etc/cni/net.d/.conffile from your CNI plugin.
4. Check for Scheduling Constraints
Taints and tolerations or node selectors might be preventing kube-dns from being scheduled.
- From the
kubectl describe podoutput, look in the Events section for messages like:0/2 nodes are available: 2 node(s) had taint {node-role.kubernetes.io/master: }, that the pod didn't tolerate.
Fix:
- If your master node has a taint that blocks scheduling (kubeadm adds this by default), you can remove it to allow Pods to run on the master:
kubectl taint nodes --all node-role.kubernetes.io/master- - Alternatively, ensure your new node is ready and doesn't have conflicting taints that kube-dns can't tolerate.
5. Check for Image Pull Failures
If the kube-dns image can't be pulled from the registry, the Pod will stay in Pending.
- In the
kubectl describe podoutput, look for events likeFailed to pull image "k8s.gcr.io/kube-dns/kube-dns:1.14.10".
Fix:
- If you have limited access to the official registry, pull the image from a mirror registry and retag it:
# Example using a public mirror docker pull registry.aliyuncs.com/google_containers/kube-dns:1.14.10 docker tag registry.aliyuncs.com/google_containers/kube-dns:1.14.10 k8s.gcr.io/kube-dns/kube-dns:1.14.10 - Restart the kube-dns Pod after fixing the image issue.
Quick First Step
Always start with describing the problematic Pod—it gives you direct insight into why scheduling is failing:
kubectl describe pod kube-dns-86f4d74b45-br6ck -n kube-system
内容的提问来源于stack exchange,提问作者user3398900

