Conjure-Up/CDK部署K8S v1.18.2后CoreDNS运行但未就绪求助
Let’s break down your issue and walk through potential fixes step by step—you’ve already done some solid troubleshooting on your own, so let’s build on that.
Background Recap
You deployed Kubernetes v1.18.2 (CDK) on Ubuntu 18.04 via conjure-up, destroyed the environment, then redeployed manually using the CDK bundle. The same DNS-related issues persist despite keeping the K8s and OS versions consistent.
Current CoreDNS Configuration
Your CoreDNS ConfigMap in the kube-system namespace is set to forward queries to /etc/resolv.conf:
Name: coredns Namespace: kube-system Labels: cdk-addons=true Annotations: Data ==== Corefile: ---- .:53 { errors health { lameduck 5s } ready kubernetes cluster.local in-addr.arpa ip6.arpa { fallthrough in-addr.arpa ip6.arpa } prometheus :9153 forward . /etc/resolv.conf cache 30 loop reload loadbalance } Events: <none>
Observed Errors & Symptoms
- CoreDNS Pods throw timeout errors trying to reach the Kubernetes API server:
E0429 09:16:42.172959 1 reflector.go:153] pkg/mod/k8s.io/client-go@v0.17.2/tools/cache/reflector.go:105: Failed to list *v1.Endpoints: Get https://10.152.183.1:443/api/v1/endpoints?limit=500&resourceVersion=0: dial tcp 10.152.183.1:443: i/o timeout [INFO] plugin/ready: Still waiting on: "kubernetes" - You can curl the Kubernetes Service endpoints successfully with the
--insecureflag, so the service itself is functional. - Most vnets created by Juju show a UNKNOWN state, and you suspect AppArmor restrictions might be contributing to this based on CDK bundle documentation.
- Attempted fixes that didn't resolve the issue:
- Modifying CoreDNS ConfigMap to use
/run/systemd/resolve/resolv.conf(config gets reverted) - Setting
kubelet-extra-config: {resolvConf: /run/systemd/resolve/resolv.conf}and rebooting (change appears in kubelet config but no fix) - Deploying Ubuntu 16.04 (Xenial) with working
/etc/resolv.conf—issue still persists
- Modifying CoreDNS ConfigMap to use
Recommended Troubleshooting Steps & Fixes
1. Prevent CoreDNS ConfigMap Reversion
Since your CoreDNS ConfigMap changes are being reverted, this is likely due to the CDK addon controller managing the CoreDNS deployment. To override this at the Juju level (so changes won’t get overwritten):
juju config coredns corefile='.:53 { errors health { lameduck 5s } ready kubernetes cluster.local in-addr.arpa ip6.arpa { fallthrough in-addr.arpa ip6.arpa } prometheus :9153 forward . /run/systemd/resolve/resolv.conf cache 30 loop reload loadbalance }'
2. Verify Kubelet DNS Configuration
Even though you set resolvConf in kubelet-extra-config, double-check consistency across all nodes:
- On each node, confirm the kubelet config file (typically
/var/lib/kubelet/config.yaml) includesresolvConf: /run/systemd/resolve/resolv.conf - Restart the kubelet service on all nodes:
sudo systemctl restart kubelet - Verify pods inherit the correct DNS settings by checking a test pod’s resolv.conf:
kubectl run -it --rm --image=ubuntu:18.04 test-pod -- bash -c "cat /etc/resolv.conf"
3. Investigate AppArmor & VNet Issues
The UNKNOWN vnet state suggests AppArmor might be blocking network traffic between CoreDNS and the API server:
- Check active AppArmor profiles:
sudo aa-status - Temporarily set the CoreDNS deployment to use an unconfined AppArmor profile to test:
kubectl patch deployment coredns -n kube-system --patch '{"spec":{"template":{"spec":{"securityContext":{"appArmorProfile":{"type":"Unconfined"}}}}}' - If this resolves the timeout, you’ll need to adjust the AppArmor profile to allow CoreDNS to communicate with the API server.
4. Validate Node Network Connectivity
Ensure all nodes can reach the Kubernetes API server IP (10.152.183.1) on port 443:
- On each node, run:
nc -zv 10.152.183.1 443 - If there’s a connection failure, check iptables rules or Calico network policies (CDK uses Calico by default) that might be blocking traffic.
5. Check CoreDNS Pod Network Namespace
Misconfigurations in the pod’s network namespace can also cause DNS issues:
- Inspect the CoreDNS pod’s network settings:
kubectl describe pod -n kube-system <coredns-pod-name> | grep -A5 "Network" - Ensure the pod is using the correct cluster network namespace and has full network access.
Final Notes
Since the issue persists across Ubuntu versions, it’s less likely to be an OS-specific resolv.conf problem and more related to cluster network policies, Juju-managed addons, or AppArmor restrictions. Start with locking in the CoreDNS config via Juju, then rule out AppArmor and network connectivity issues one by one.
内容的提问来源于stack exchange,提问作者Dom

