You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes Pod无法创建或终止问题排查求助

Hey there, let’s break down your problem step by step and work through the troubleshooting process.

Absolutely. Your error message clearly states that the CNI plugin Calico can’t connect to the API Server at https://[10.96.0.1]:443 to fetch ClusterInformation, with a "no route to host" warning. This is the root cause of most of your problems:

  • Pods can’t be created because their sandbox can’t initialize without a working CNI plugin.
  • Terminating Pods get stuck because the kubelet or CNI can’t communicate with the API Server to complete cleanup.
  • The flaky state of calico-node, coredns, and metrics-server Pods all ties back to this network connectivity failure between Calico and the API Server.

How to deep-dive into the "server refused the request for unknown reasons" error

Let’s start with targeted checks to pinpoint why Calico can’t reach the API Server:

  1. Check the problematic calico-node Pod logs
    Run this command to pull logs from the failing calico-node Pod:

    kubectl logs -n kube-system <your-calico-node-pod-name>
    

    Look for keywords like:

    • connection refused or timeout when reaching 10.96.0.1:443
    • Certificate validation errors (indicates misconfigured CA certs)
    • BGP neighbor failures (if Calico uses BGP mode)
    • IP address conflicts or missing route entries
  2. Verify Calico’s API Server configuration
    Inspect the calico-node Pod’s environment variables and mounts to ensure it can reach the API Server:

    kubectl describe pod -n kube-system <your-calico-node-pod-name>
    
    • Check if KUBERNETES_SERVICE_HOST and KUBERNETES_SERVICE_PORT are set to 10.96.0.1 and 443 respectively.
    • Confirm the cluster CA cert is mounted correctly (look for volumes pointing to /etc/kubernetes/pki/ca.crt and corresponding volume mounts in the Pod).
  3. Test connectivity directly from the node
    Log into the node with the failing calico-node Pod, then run:

    # Test API Server reachability with cluster CA cert
    curl -v https://10.96.0.1:443 --cacert /etc/kubernetes/pki/ca.crt
    # Check if there's a route to the ClusterIP subnet (10.96.0.0/12 by default)
    ip route show | grep 10.96.0.0
    
    • If the curl command fails with "no route", the node lacks a route to the API Server’s ClusterIP. This is usually managed by kube-proxy, so check if kube-proxy Pods are running normally.
    • If the route exists but curl still fails, inspect iptables rules for API Server traffic:
      iptables-save | grep KUBE-SVC-KUBEAPI
      
      Missing rules here mean kube-proxy isn’t maintaining the necessary forwarding rules.

Troubleshooting stuck Terminating Pods

For Pods that show as Terminating but throw "Pod does not exist" when described:

  1. Force clean up stuck Pods
    Run this command to bypass graceful termination (note: only do this after addressing the root Calico issue, otherwise new Pods will get stuck too):
    kubectl delete pod <pod-name> -n <namespace> --force --grace-period=0
    
  2. Check kubelet logs on the node
    The kubelet handles Pod cleanup, so look for errors here:
    journalctl -u kubelet -f
    
    Look for messages about failed container runtime calls (docker/containerd) or network plugin timeouts.

Next steps for further investigation

  1. Validate Pod CIDR consistency
    Kubeadm and Calico must use matching Pod CIDRs. Check:

    • Kubeadm’s configured Pod CIDR:
      kubeadm config view | grep podSubnet
      
    • Calico’s configured Pod CIDR:
      kubectl get configmap -n kube-system calico-config -o yaml | grep CALICO_IPV4POOL_CIDR
      

    If these don’t match (e.g., kubeadm uses 10.244.0.0/16 and Calico uses 192.168.0.0/16), Calico can’t set up proper routing, leading to API Server connectivity issues.

  2. Check Calico node-to-node communication
    If Calico uses BGP mode (default), verify neighbor connectivity:

    • Install calicoctl on a node, then run:
      calicoctl node status
      

    Look for failed BGP peers. If you’re in a cloud environment, ensure security groups open port 179 (BGP) and 4789 (IPIP tunnel, if used).

  3. Roll back recent Helm changes
    Since the issue started after using Helm, roll back your last Helm release to rule out conflicts:

    # List all releases to find the problematic one
    helm list -A
    # Roll back to a previous working revision
    helm rollback <release-name> <revision-number>
    
  4. Recheck kubelet and kube-proxy status
    Ensure both services are running and healthy on all nodes:

    systemctl status kubelet
    kubectl get pods -n kube-system | grep kube-proxy
    

    Restart kubelet if needed: systemctl restart kubelet

内容的提问来源于stack exchange,提问作者Leandro De Mestico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:57:33