You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复Kubernetes节点NetworkUnavailable:True错误并排查根源

Hey there, let's walk through how to troubleshoot that frustrating Kubernetes node network issue where you saw Status NetworkUnavailable: True paired with the error RouteController failed to create a route. Even though deleting and recreating the node fixed it temporarily, we need to get to the bottom of why this happened so you can prevent it from popping up again.

Troubleshooting Kubernetes Node NetworkUnavailable (RouteController Route Creation Failure)

1. Start with the Route Controller Logs

First things first—let's look at what the route controller itself is complaining about. The route controller lives in the kube-system namespace (the exact pod name depends on your CNI: Calico, Flannel, your cloud provider's CNI, etc.). Grab its logs with:
kubectl logs -n kube-system <your-route-controller-pod-name>
Scan for specific errors around route creation: maybe it's a permission issue (like not being able to modify VPC routes on AWS), overlapping pod CIDRs, or it can't reach the node's network interface. These logs will give you the most direct clue about what went wrong.

2. Verify Node Network Configuration

Check if the node's network setup is as expected:

  • List the node's network interfaces and IPs to confirm the primary interface has the correct address:
    ip addr show
  • Ensure the node's assigned PodCIDR matches what the cluster expects:
    kubectl describe node <problematic-node-name> | grep PodCIDR
    If the PodCIDR is missing or incorrect, the route controller can't create a valid route to direct pod traffic to this node.

3. Inspect Your CNI Plugin Health

Different CNIs handle routing differently—let's check the CNI components on the problematic node:

  • For Calico: Check the calico-node pod logs on the node for BGP peering failures or route sync issues:
    kubectl logs -n kube-system <calico-node-pod-name> -c calico-node
  • For cloud provider CNIs (like AWS VPC CNI): Verify the node has the required cluster tags (e.g., kubernetes.io/cluster/<cluster-name>: owned)—the route controller uses these to identify which nodes need routes.
  • For Flannel: Check if the flanneld pod on the node is running and has no errors in its logs:
    kubectl logs -n kube-system <flannel-pod-name>

4. Check Network Infrastructure Limits

On cloud platforms or on-prem clusters, infrastructure constraints often cause route creation failures:

  • Cloud environments: Most providers have limits on the number of routes per VPC subnet. Check your cloud console to see if you've hit the route table limit—if so, you'll need to clean up unused routes or request a limit increase.
  • On-prem networks: Verify that switches/routers allow route propagation from the Kubernetes controller. ACLs or firewall rules might be blocking the controller from adding new routes to the network.

5. Validate Kubelet and Node Registration

A misbehaving kubelet or incomplete node registration can break route setup:

  • Describe the node to check kubelet health and registration status:
    kubectl describe node <problematic-node-name>
    Look for conditions like Ready status and any kubelet error messages (e.g., failed to report node IP to the API server).
  • Ensure the kubelet is using the correct --node-ip flag if your node has multiple network interfaces—an incorrect node IP will make the route controller point traffic to the wrong address.

6. Look for Race Conditions or Temporary Glitches

Since recreating the node fixed the issue, it might have been a one-time bootstrap glitch:

  • Check the node's events around the time of the failure to see if there were network provisioning delays:
    kubectl get events --field-selector involvedObject.name=<problematic-node-name> --sort-by='.metadata.creationTimestamp'
    Sometimes cloud providers don't fully set up the node's network before the kubelet starts, leading to route creation failures that resolve on a restart.

7. Check Version Compatibility

Outdated Kubernetes or CNI versions can have known routing bugs:

  • Verify that your CNI plugin version is compatible with your cluster's Kubernetes version (check the CNI's official docs for compatibility matrices).
  • If you're running older versions, consider upgrading the CNI plugin or cluster to patch known issues with route controller logic.

内容的提问来源于stack exchange,提问作者jirawat paiboon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:49:40