KubeDNS异常及节点DNS解析失败问题排查求助
Let's break down your issues step by step—you've got two core problems here: node hostname resolution failures blocking kubectl exec from reaching the worker node, and kube-dns failing to communicate with the Kubernetes API server due to routing gaps. Here's how to diagnose and fix them:
1. Fix Node Hostname Resolution (the lookup worker2 on 127.0.0.53:53: server misbehaving error)
That error means your Kubernetes API server (running on master nodes) can't resolve the worker2 hostname to its internal IP when trying to connect to the node's kubelet. Here's what to check:
Verify
/etc/hostson all nodes: Since you're using Kubernetes The Hard Way, manual hostfile entries are usually required for cluster nodes. On every master and worker node, open/etc/hostsand ensure you have lines for every node like this:10.133.55.62 kube3 10.133.52.77 kube1 10.133.55.73 kube2 10.133.56.88 worker1 10.133.55.220 worker3 10.133.56.89 worker2DigitalOcean's internal DNS doesn't auto-resolve custom droplet hostnames, so these manual entries are critical.
Test resolution on master nodes: SSH into a master node (e.g., kube3) and run
nslookup worker2orping worker2. You should see the internal IP10.133.56.89returned immediately. If it resolves to a public IP or fails, your hostfile is missing entries or there's a systemd-resolved conflict.
2. Diagnose Flannel Network & kube-dns Routing Issues
The kube-dns error about no route to 10.32.0.1:443 (the default Kubernetes Service IP for the API server) points to a broken Pod-to-Service network path. Let's dig into this:
Check Flannel Health
- Verify Flannel interface configuration: On every node, run
ip addr show flannel.1. You should see an IP in the10.244.0.0/16range, with no overlapping subnets across nodes. If the interface doesn't exist, Flannel isn't running properly. - Check flanneld service status: Run
systemctl status flanneld(assuming systemd deployment). Look for errors like failed network config fetch or API server connection refusals. Restart the service if needed withsystemctl restart flanneld. - Test cross-node internal connectivity: From any master node, ping worker2's internal IP (
10.133.56.89). If this fails, check DigitalOcean's cloud firewall rules—ensure all internal traffic (especially UDP port 8472 for Flannel VXLAN) is allowed between cluster droplets.
Validate Service Network (10.32.0.0/24) Routing
- Test API server access from kube-dns node: First, find which node runs your kube-dns Pod:
SSH into that node, then runkubectl get pods -n kube-system -o wide | grep kube-dnscurl -k https://10.32.0.1:443. You should get a 403 Forbidden response (normal, since you're hitting the API server without auth). If you get a "no route to host" error, kube-proxy isn't setting up necessary iptables rules. - Check kube-proxy iptables rules: On the same node, run:
You should see multiple rules forwarding traffic toiptables-save | grep 10.32.0.110.32.0.1to your master nodes' API server IPs. If no rules exist, restart kube-proxy withsystemctl restart kube-proxyand check its logs for sync errors. - Verify kube-dns config: Ensure kube-dns points to the correct API server address:
Look for thekubectl get configmap kube-dns -n kube-system -o yamlkubeMasterUrlfield under thekubednssection—it should be set tohttps://10.32.0.1:443(matching your Service cluster network).
3. Cross-Verify Pod-to-Pod Connectivity
To rule out broader network issues, pick a Pod IP on worker1 and ping it from worker2 (or vice versa). If this fails, Flannel's VXLAN tunnel isn't working—double-check that UDP 8472 is open on all nodes' firewalls, and that Flannel's network config matches your 10.244.0.0/16 Pod CIDR.
内容的提问来源于stack exchange,提问作者Richard87

