Kubernetes v1.10集群kube-dns CrashLoopBackOff问题求助
Hey there, let's work through this kube-dns problem you're hitting with your Kubernetes 1.10 cluster running flannel. That Failed to list *v1.Endpoints getsockopt: connection refused error and the endless waiting logs tell us kube-dns can't reach the Kubernetes API server—this is the root of your CrashLoopBackOff issue. Here's how to diagnose and fix it step by step:
1. Test API Server Connectivity Directly from the kube-dns Pod
First, confirm the kube-dns pod can actually reach the API server. Grab the name of your kube-dns pod first:
kubectl get pods -n kube-system
Then exec into the pod and test connectivity to the API server's default service IP (usually 10.96.0.1):
kubectl exec -it <your-kube-dns-pod-name> -n kube-system sh # Inside the pod, run this to check connectivity curl -v http://10.96.0.1:8080
If you get a connection refused here, that confirms the API server is unreachable from the pod's network.
2. Validate Flannel Network & API Server Listening Address
Flannel relies on the API server being accessible to pods, so let's check key network configurations:
- Check API Server's bind address: Ensure the kube-apiserver isn't only listening on localhost (which pods can't reach). Inspect the apiserver pod's arguments:
Look forkubectl describe pod kube-apiserver-<your-node-name> -n kube-system | grep -A 20 "Args:"--insecure-bind-address=0.0.0.0or--bind-address=<cluster-internal-ip>(not127.0.0.1). If it's stuck on localhost, update the apiserver config to listen on a cluster-accessible IP. - Match Flannel CIDR with API Server: Verify flannel's pod network CIDR matches what's configured in the API server. Check the API server's
--pod-network-cidrarg, then compare it to flannel's config:
Thekubectl get configmap kube-flannel-cfg -n kube-system -o yamlNetworkvalue in flannel's config should exactly match the API server's pod CIDR.
3. Ensure kube-dns Has Proper RBAC Permissions
Kubernetes v1.10 uses RBAC by default, so kube-dns needs the right permissions to list services and endpoints:
- Check if the kube-dns service account exists:
kubectl get serviceaccount kube-dns -n kube-system - Verify the cluster role binding for kube-dns:
If the binding is missing, create it with this manifest:kubectl get clusterrolebinding kube-dns
Apply it withapiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRoleBinding metadata: name: kube-dns roleRef: apiGroup: rbac.authorization.k8s.io kind: ClusterRole name: system:kube-dns subjects: - kind: ServiceAccount name: kube-dns namespace: kube-systemkubectl apply -f <your-file-name>.yaml.
4. Check kube-dns Deployment Configuration
Make sure the kube-dns deployment is pointing to the correct API server:
kubectl describe deployment kube-dns -n kube-system | grep -A 10 "Args:"
Look for --kube-master-url=http://<api-server-ip>:8080 or confirm it's using the default API service IP (10.96.0.1). If the URL is incorrect, update the deployment to fix it.
5. Verify Node & Flannel Health
Ensure flannel is running properly across all nodes and no firewalls are blocking traffic:
- Check flannel pods are in
Runningstate:kubectl get pods -n kube-system | grep flannel - On each node, confirm the flannel interface exists and has a valid IP:
ip addr show flannel.1 - Double-check node firewalls aren't blocking port 8080 (insecure API port) or 6443 (secure API port) from the pod network.
Final Step: Restart kube-dns Pods
After fixing any issues above, delete the existing kube-dns pods to force a restart:
kubectl delete pods -n kube-system -l k8s-app=kube-dns
Then check the logs again with kubectl logs <new-kube-dns-pod> -n kube-system—the waiting messages should stop, and kube-dns should initialize successfully.
内容的提问来源于stack exchange,提问作者junglie85

