Kubernetes API Server启动失败:无法连接DNS服务器求助
Hey Duncan, sorry to hear your cluster went offline out of the blue—let’s break down what’s happening here and work through practical fixes to get things back up.
First, let’s zero in on the fatal error that’s stopping your API Server from starting up; this is the root cause driving the other secondary errors:
F0609 00:03:14.534558 1 controller.go:161] Unable to perform initial IP allocation check: unable to refresh the service IP block: Get https://localhost:6443/api/v1/services : dial tcp: lookup localhost on 172.16.0.1:53: no such host
The earlier error about failing to remove old endpoints is just a side effect of the API Server not being able to connect to itself. Here’s what to check and fix step by step:
1. Dig into the DNS Server the API Server is Using
You mentioned your local DNS is still running fine, but the API Server is trying to use 172.16.0.1:53 to resolve localhost—this is odd, because localhost should resolve locally without needing to hit an external DNS server.
- Check the API Server pod’s DNS config:
Runkubectl describe pod kube-apiserver-<your-node-name> -n kube-systemand look for theDNS Configsection.172.16.0.1is likely your cluster’s CoreDNS service IP—if CoreDNS is misconfigured or crashed, it won’t handlelocalhostqueries correctly. - Test the DNS server manually on the control plane node:
If this command fails, that confirms the DNS server atdig @172.16.0.1 localhost172.16.0.1isn’t resolvinglocalhostproperly.
2. Bypass DNS to Get the API Server Up Quickly
To restore cluster access fast, modify the API Server’s config to use an IP directly instead of localhost, so it doesn’t need DNS resolution:
- Edit the kube-apiserver static pod manifest (usually located at
/etc/kubernetes/manifests/kube-apiserver.yamlon your control plane node). - Find any references to
https://localhost:6443and replacelocalhostwith127.0.0.1or your control plane node’s public/private IP. - Also, make sure the
--bind-addressflag is set to127.0.0.1or your node’s IP to ensure the API Server listens on a resolvable address. - Save the file—kubelet will automatically restart the API Server pod with the updated config.
3. Check Node-Level DNS Settings
Even if you didn’t touch cluster config, node network settings might have changed unexpectedly:
- Look at your control plane node’s
/etc/resolv.conffile—if it listsnameserver 172.16.0.1, this could be propagating to the API Server pod. - Verify kubelet’s DNS config: Check the kubelet config file (e.g.,
/var/lib/kubelet/config.yaml) or systemd service file forclusterDNSorresolvConfentries that might be overriding DNS settings.
4. Fix the Etcd Storage Error (Once the API Server is Up)
After getting the API Server running, address the secondary storage error:
E0609 00:03:14.518792 1 controller.go:152] Unable to remove old endpoints from kubernetes service: StorageError: key not found, Code: 1, Key: /registry/masterleases/192.168.5.2, ResourceVersion: 0, AdditionalErrorMsg:
- First, confirm etcd is healthy: Run
etcdctl endpoint healthto check connectivity and status. - If etcd is working, you can clean up the stale master lease key (make sure to back up etcd first!):
etcdctl del /registry/masterleases/192.168.5.2
Start with step 2 to get your cluster back online, then work through the other steps to fix the root cause. Let me know if you hit any roadblocks along the way!
内容的提问来源于stack exchange,提问作者Duncan Krebs

