Kubernetes跨节点Pod IP无法连通问题求助
Alright, let's dig into this cross-node pod connectivity issue you're dealing with. From what you've shared, pods are only reachable from the node they're running on—cross-node pings (like hitting the nav pod 192.168.104.8 on node2 from node3) are dropping 100% packets. Let's break down the problem and walk through the fixes step by step.
- Core Symptom: Pod IPs respond to pings only from their host node; any cross-node ping attempt results in full packet loss. For example:
Pinging
192.168.104.8(nav-6f67d5bd79-9khmm pod on node2) from node3 returns 100% packet loss.
Environment Details
kube-system Core & Calico Pods
master2@master2:~$ kubectl get pods --namespace=kube-system -o wide NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES calico-kube-controllers-6ff8cbb789-lxwqq 1/1 Running 0 6d21h 192.168.180.2 master2 <none> <none> calico-node-4mnfk 1/1 Running 0 4d20h 10.10.41.165 node3 <none> <none> calico-node-c4rjb 1/1 Running 0 6d21h 10.10.41.159 master2 <none> <none> calico-node-dgqwx 1/1 Running 0 4d20h 10.10.41.153 master1 <none> <none> calico-node-fhtvz 1/1 Running 0 6d21h 10.10.41.161 node2 <none> <none> calico-node-mhd7w 1/1 Running 0 4d21h 10.10.41.155 node1 <none> <none> coredns-8b5d5b85f-fjq72 1/1 Running 0 45m 192.168.135.11 node3 <none> <none> coredns-8b5d5b85f-hgg94 1/1 Running 0 45m 192.168.166.136 node1 <none> <none> etcd-master1 1/1 Running 0 4d20h 10.10.41.153 master1 <none> <none> etcd-master2 1/1 Running 0 6d21h 10.10.41.159 master2 <none> <none> kube-apiserver-master1 1/1 Running 0 4d20h 10.10.41.153 master1 <none> <none> kube-apiserver-master2 1/1 Running 0 6d21h 10.10.41.159 master2 <none> <none> kube-controller-manager-master1 1/1 Running 0 4d20h 10.10.41.153 master1 <none> <none> kube-controller-manager-master2 1/1 Running 2 6d21h 10.10.41.159 master2 <none> <none> kube-proxy-66nxz 1/1 Running 0 6d21h 10.10.41.159 master2 <none> <none> kube-proxy-fnrrz 1/1 Running 0 4d20h 10.10.41.153 master1 <none> <none> kube-proxy-lq5xp 1/1 Running 0 6d21h 10.10.41.161 node2 <none> <none> kube-proxy-vxhwm 1/1 Running 0 4d21h 10.10.41.155 node1 <none> <none> kube-proxy-zgwzq 1/1 Running 0 4d20h 10.10.41.165 node3 <none> <none> kube-scheduler-master1 1/1 Running 0 4d20h 10.10.41.153 master1 <none> <none> kube-scheduler-master2 1/1 Running 1 6d21h 10.10.41.159 master2 <none> <none>
Application Pods
master1@master1:~/cluster$ sudo kubectl get pods -o wide NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES contentms-cb475f569-t54c2 1/1 Running 0 6d21h 192.168.104.1 node2 <none> <none> nav-6f67d5bd79-9khmm 1/1 Running 0 6d8h 192.168.104.8 node2 <none> <none> react 1/1 Running 0 7m24s 192.168.135.12 node3 <none> <none> statistics-5668cd7dd-thqdf 1/1 Running 0 6d15h 192.168.104.4 node2 <none> <none>
Since you're using Calico for networking, the issue almost always ties to BGP peerings, route advertisements, or firewall rules blocking cross-node traffic. Let's start with the most likely fixes:
1. Validate Calico BGP Peer Connections
Calico relies on BGP to share pod route information between nodes. If peers aren't established, cross-node traffic can't flow.
Run this on any master node to check peer status (replace calico-node-4mnfk with a calico pod from your cluster):
kubectl exec -n kube-system calico-node-4mnfk -- calicoctl node status
You should see all cluster nodes listed as peers with an Established state. If any peer is missing or stuck in Idle/Connect, BGP isn't forming.
Fix: Ensure all nodes allow TCP traffic on port 179 (BGP's default port). On Ubuntu nodes, you can run:
sudo ufw allow 179/tcp sudo ufw reload
For other distros, adjust the firewall command to match your system (e.g., firewalld on RHEL/CentOS).
2. Check Cross-Node Route Advertisements
Calico should advertise pod CIDR routes to all other nodes. Let's verify node3 has a route to node2's pod CIDR (192.168.104.0/24):
On node3, run:
ip route show | grep 192.168.104
You should see a route pointing to node2's IP (10.10.41.161). If no route exists, Calico isn't advertising it correctly.
Fix: Restart the Calico pods on node2 and node3 to refresh BGP advertisements:
kubectl delete pod -n kube-system calico-node-fhtvz calico-node-4mnfk
Kubernetes will automatically recreate these pods.
3. Verify Calico IP Pool Configuration
Make sure Calico's IP pool includes all your pod CIDR ranges. Run:
kubectl exec -n kube-system calico-node-4mnfk -- calicoctl get ippools -o yaml
Check that the cidr field covers ranges like 192.168.104.0/24 and 192.168.135.0/24. If not, update the IP pool to include these ranges.
4. Check Node Firewall Rules for Calico Traffic
Even if BGP works, node-level iptables rules might block pod-to-pod traffic. On node3, verify Calico's rules exist:
sudo iptables -L -n | grep CALICO
You should see rules allowing traffic from other pod CIDRs. If missing, restart the Calico node pod to regenerate iptables rules.
After applying the fixes, test the ping again from node3 to the nav pod:
ping 192.168.104.8
You should start seeing successful responses if the issue is resolved.
内容的提问来源于stack exchange,提问作者piyush

