Kubernetes集群Pod到Service网络连接问题排查(WeaveNet、kube-dns)
Hey there! Let's dig into your kube-dns issue that's blocking your Network Policies testing and ELK Stack log export. Since you mentioned name resolution seems to work but Pod-to-Service connections are acting up, here's a step-by-step troubleshooting plan tailored to your 3-node Kubernetes cluster with WeaveNet:
First, let's make sure the DNS components themselves are running properly:
- Run
kubectl get pods -n kube-system -l k8s-app=kube-dnsto check if all DNS pods are in aRunningstate with full readiness. If any are crashing or not ready, runkubectl describe pod <pod-name> -n kube-systemto spot issues like image pull failures, resource constraints, or network misconfigs. - Check the DNS pod logs with
kubectl logs <coredns-pod-name> -n kube-system -c coredns(note: most modern clusters use CoreDNS instead of the old kube-dns, so adjust the label if needed). Look for errors like "connection refused" to backend services, or warnings about WeaveNet integration issues. - Confirm the DNS service has valid endpoints: Run
kubectl get svc kube-dns -n kube-systemto get the service IP, thenkubectl get endpoints kube-dns -n kube-system— the endpoints should match the IPs of your running CoreDNS pods.
Even if nslookup works, there might be a disconnect between resolution and real connectivity:
- Spin up a test pod with diagnostic tools:
kubectl run test-pod --image=busybox:1.28 --rm -it -- sh - Inside the test pod, first confirm DNS resolution:
nslookup kubernetes.default.svc.cluster.local— you should get the cluster IP of the default kubernetes service. - Then test direct connectivity to the service: Use
wget -qO- http://kubernetes.default.svc.cluster.local:443ortelnet kubernetes.default.svc.cluster.local 443. If this fails, check if the target service's endpoints are healthy, or if a Network Policy is accidentally blocking traffic.
Since you're using WeaveNet, let's rule out issues with how it interacts with Network Policies:
- Ensure WeaveNet pods are running on every node:
kubectl get pods -n kube-system -l name=weave-net— each node should have exactly one WeaveNet pod inRunningstate. - Confirm Network Policies are enabled: On your control plane node, run
ps aux | grep kube-apiserverand check for the flag--enable-admission-plugins=...,NetworkPolicy(it should include NetworkPolicy in the list). WeaveNet supports Network Policies, but this admission plugin needs to be enabled for them to work. - Temporarily disable all existing Network Policies to test: Run
kubectl delete networkpolicy --all(make sure you have backups to restore them later!). If connectivity starts working after this, one of your policies is blocking DNS or service traffic.
Double-check the DNS settings that pods inherit:
- Look at a problematic pod's DNS config:
kubectl describe pod <your-app-pod> | grep -A5 DNS— it should list the kube-dns service IP as the nameserver, and the search domain should includesvc.cluster.local. - Verify kubelet DNS settings on each node: Check
/var/lib/kubelet/config.yamlor the kubelet command line flags for--cluster-dnsand--cluster-domain. These should point to the kube-dns service IP andcluster.localrespectively.
Once you resolve the kube-dns issue, here's a quick check to ensure log export to ELK goes smoothly:
- Deploy a log collector like Fluentd or Filebeat as a DaemonSet. Make sure the collector pods can resolve your Elasticsearch service name (if ELK is running in-cluster) or external endpoint.
- Update your Network Policies to allow outbound traffic from the log collector pods to Elasticsearch and Kibana ports (typically 9200 for Elasticsearch, 5601 for Kibana).
If you can share specific error messages from your failed Pod-to-Service connections, or the output of any of the commands above, we can narrow this down even further!
内容的提问来源于stack exchange,提问作者Verena I.

