You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes集群Pod到Service网络连接问题排查(WeaveNet、kube-dns)

Hey there! Let's dig into your kube-dns issue that's blocking your Network Policies testing and ELK Stack log export. Since you mentioned name resolution seems to work but Pod-to-Service connections are acting up, here's a step-by-step troubleshooting plan tailored to your 3-node Kubernetes cluster with WeaveNet:

1. Verify kube-dns/CoreDNS Core Health

First, let's make sure the DNS components themselves are running properly:

  • Run kubectl get pods -n kube-system -l k8s-app=kube-dns to check if all DNS pods are in a Running state with full readiness. If any are crashing or not ready, run kubectl describe pod <pod-name> -n kube-system to spot issues like image pull failures, resource constraints, or network misconfigs.
  • Check the DNS pod logs with kubectl logs <coredns-pod-name> -n kube-system -c coredns (note: most modern clusters use CoreDNS instead of the old kube-dns, so adjust the label if needed). Look for errors like "connection refused" to backend services, or warnings about WeaveNet integration issues.
  • Confirm the DNS service has valid endpoints: Run kubectl get svc kube-dns -n kube-system to get the service IP, then kubectl get endpoints kube-dns -n kube-system — the endpoints should match the IPs of your running CoreDNS pods.
2. Validate Pod-to-Service Resolution & Actual Connectivity

Even if nslookup works, there might be a disconnect between resolution and real connectivity:

  • Spin up a test pod with diagnostic tools: kubectl run test-pod --image=busybox:1.28 --rm -it -- sh
  • Inside the test pod, first confirm DNS resolution: nslookup kubernetes.default.svc.cluster.local — you should get the cluster IP of the default kubernetes service.
  • Then test direct connectivity to the service: Use wget -qO- http://kubernetes.default.svc.cluster.local:443 or telnet kubernetes.default.svc.cluster.local 443. If this fails, check if the target service's endpoints are healthy, or if a Network Policy is accidentally blocking traffic.
3. Check WeaveNet & Network Policy Integration

Since you're using WeaveNet, let's rule out issues with how it interacts with Network Policies:

  • Ensure WeaveNet pods are running on every node: kubectl get pods -n kube-system -l name=weave-net — each node should have exactly one WeaveNet pod in Running state.
  • Confirm Network Policies are enabled: On your control plane node, run ps aux | grep kube-apiserver and check for the flag --enable-admission-plugins=...,NetworkPolicy (it should include NetworkPolicy in the list). WeaveNet supports Network Policies, but this admission plugin needs to be enabled for them to work.
  • Temporarily disable all existing Network Policies to test: Run kubectl delete networkpolicy --all (make sure you have backups to restore them later!). If connectivity starts working after this, one of your policies is blocking DNS or service traffic.
4. Inspect Cluster-Wide DNS Configuration

Double-check the DNS settings that pods inherit:

  • Look at a problematic pod's DNS config: kubectl describe pod <your-app-pod> | grep -A5 DNS — it should list the kube-dns service IP as the nameserver, and the search domain should include svc.cluster.local.
  • Verify kubelet DNS settings on each node: Check /var/lib/kubelet/config.yaml or the kubelet command line flags for --cluster-dns and --cluster-domain. These should point to the kube-dns service IP and cluster.local respectively.
5. Quick Prep for ELK Log Export (Post-DNS Fix)

Once you resolve the kube-dns issue, here's a quick check to ensure log export to ELK goes smoothly:

  • Deploy a log collector like Fluentd or Filebeat as a DaemonSet. Make sure the collector pods can resolve your Elasticsearch service name (if ELK is running in-cluster) or external endpoint.
  • Update your Network Policies to allow outbound traffic from the log collector pods to Elasticsearch and Kibana ports (typically 9200 for Elasticsearch, 5601 for Kibana).

If you can share specific error messages from your failed Pod-to-Service connections, or the output of any of the commands above, we can narrow this down even further!

内容的提问来源于stack exchange,提问作者Verena I.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:20:57