Kubernetes Pod无法创建或终止问题排查求助
Hey there, let’s break down your problem step by step and work through the troubleshooting process.
First: Is this issue related to Calico?
Absolutely. Your error message clearly states that the CNI plugin Calico can’t connect to the API Server at https://[10.96.0.1]:443 to fetch ClusterInformation, with a "no route to host" warning. This is the root cause of most of your problems:
- Pods can’t be created because their sandbox can’t initialize without a working CNI plugin.
- Terminating Pods get stuck because the kubelet or CNI can’t communicate with the API Server to complete cleanup.
- The flaky state of
calico-node,coredns, andmetrics-serverPods all ties back to this network connectivity failure between Calico and the API Server.
How to deep-dive into the "server refused the request for unknown reasons" error
Let’s start with targeted checks to pinpoint why Calico can’t reach the API Server:
Check the problematic calico-node Pod logs
Run this command to pull logs from the failingcalico-nodePod:kubectl logs -n kube-system <your-calico-node-pod-name>Look for keywords like:
connection refusedortimeoutwhen reaching 10.96.0.1:443- Certificate validation errors (indicates misconfigured CA certs)
- BGP neighbor failures (if Calico uses BGP mode)
- IP address conflicts or missing route entries
Verify Calico’s API Server configuration
Inspect thecalico-nodePod’s environment variables and mounts to ensure it can reach the API Server:kubectl describe pod -n kube-system <your-calico-node-pod-name>- Check if
KUBERNETES_SERVICE_HOSTandKUBERNETES_SERVICE_PORTare set to10.96.0.1and443respectively. - Confirm the cluster CA cert is mounted correctly (look for volumes pointing to
/etc/kubernetes/pki/ca.crtand corresponding volume mounts in the Pod).
- Check if
Test connectivity directly from the node
Log into the node with the failingcalico-nodePod, then run:# Test API Server reachability with cluster CA cert curl -v https://10.96.0.1:443 --cacert /etc/kubernetes/pki/ca.crt # Check if there's a route to the ClusterIP subnet (10.96.0.0/12 by default) ip route show | grep 10.96.0.0- If the curl command fails with "no route", the node lacks a route to the API Server’s ClusterIP. This is usually managed by kube-proxy, so check if
kube-proxyPods are running normally. - If the route exists but curl still fails, inspect iptables rules for API Server traffic:
Missing rules here mean kube-proxy isn’t maintaining the necessary forwarding rules.iptables-save | grep KUBE-SVC-KUBEAPI
- If the curl command fails with "no route", the node lacks a route to the API Server’s ClusterIP. This is usually managed by kube-proxy, so check if
Troubleshooting stuck Terminating Pods
For Pods that show as Terminating but throw "Pod does not exist" when described:
- Force clean up stuck Pods
Run this command to bypass graceful termination (note: only do this after addressing the root Calico issue, otherwise new Pods will get stuck too):kubectl delete pod <pod-name> -n <namespace> --force --grace-period=0 - Check kubelet logs on the node
The kubelet handles Pod cleanup, so look for errors here:
Look for messages about failed container runtime calls (docker/containerd) or network plugin timeouts.journalctl -u kubelet -f
Next steps for further investigation
Validate Pod CIDR consistency
Kubeadm and Calico must use matching Pod CIDRs. Check:- Kubeadm’s configured Pod CIDR:
kubeadm config view | grep podSubnet - Calico’s configured Pod CIDR:
kubectl get configmap -n kube-system calico-config -o yaml | grep CALICO_IPV4POOL_CIDR
If these don’t match (e.g., kubeadm uses
10.244.0.0/16and Calico uses192.168.0.0/16), Calico can’t set up proper routing, leading to API Server connectivity issues.- Kubeadm’s configured Pod CIDR:
Check Calico node-to-node communication
If Calico uses BGP mode (default), verify neighbor connectivity:- Install
calicoctlon a node, then run:calicoctl node status
Look for failed BGP peers. If you’re in a cloud environment, ensure security groups open port 179 (BGP) and 4789 (IPIP tunnel, if used).
- Install
Roll back recent Helm changes
Since the issue started after using Helm, roll back your last Helm release to rule out conflicts:# List all releases to find the problematic one helm list -A # Roll back to a previous working revision helm rollback <release-name> <revision-number>Recheck kubelet and kube-proxy status
Ensure both services are running and healthy on all nodes:systemctl status kubelet kubectl get pods -n kube-system | grep kube-proxyRestart kubelet if needed:
systemctl restart kubelet
内容的提问来源于stack exchange,提问作者Leandro De Mestico

