多Master高可用Kubernetes集群Flannel组件CrashLoopBackOff故障求助
First off, let's start by diagnosing the root causes behind your Flannel and CoreDNS issues—since CoreDNS is stuck in ContainerCreating, that's almost certainly because the pod network (Flannel) isn't functioning properly. Let's break down the steps to resolve this:
1. Verify Pod CIDR Alignment Between Kubeadm and Flannel
Flannel defaults to using the 10.244.0.0/16 Pod CIDR, but looking at your config.yaml, you didn't specify this in your ClusterConfiguration. Kubernetes needs this to assign pod IPs correctly.
- Check if the kube-controller-manager has the correct
--cluster-cidrflag:kubectl get pods kube-controller-manager-master1 -n kube-system -o yaml | grep cluster-cidr - If the flag is missing, edit the static pod manifest on each master node:
Add this line under thesudo vi /etc/kubernetes/manifests/kube-controller-manager.yamlcommandsection:
The kubelet will automatically restart the controller-manager pod once you save the file.- --cluster-cidr=10.244.0.0/16
2. Fix Flannel's Etcd Configuration (Critical for External Etcd)
Your external etcd cluster has client certificate authentication enabled, but the default Flannel manifest doesn't account for this. Flannel needs valid certificates to communicate with your etcd cluster, and it's currently trying to use the default local etcd endpoint (http://127.0.0.1:2379) which won't work.
Step 2.1: Customize the Flannel Manifest
- Download the original Flannel YAML:
wget https://raw.githubusercontent.com/coreos/flannel/master/Documentation/kube-flannel.yml - Edit the
kube-flannel-cfgConfigMap:
Find thenet-conf.jsonsection and update it to include your etcd endpoints and certificate paths:{ "Network": "10.244.0.0/16", "Backend": { "Type": "vxlan" }, "EtcdEndpoints": "https://192.168.1.21:2379,https://192.168.1.22:2379,https://192.168.1.23:2379", "EtcdCAFile": "/etc/etcd/ca.pem", "EtcdCertFile": "/etc/etcd/kubernetes.pem", "EtcdKeyFile": "/etc/etcd/kubernetes-key.pem" } - Add a volume mount for etcd certificates in the Flannel DaemonSet:
- Under the
volumessection of the DaemonSet, add:- name: etcd-certs hostPath: path: /etc/etcd type: Directory - Under the
containers[0].volumeMountssection, add:- name: etcd-certs mountPath: /etc/etcd readOnly: true
- Under the
Step 2.2: Reapply the Modified Flannel Manifest
- Delete the existing Flannel resources:
kubectl delete -f kube-flannel.yml - Apply the customized manifest:
kubectl apply -f kube-flannel.yml
3. Validate Network Connectivity Between Nodes
Flannel uses VXLAN (UDP port 8472) for pod-to-pod communication, and your nodes need access to the etcd ports (TCP 2379/2380).
- Test connectivity to etcd from a worker node:
ETCDCTL_API=3 etcdctl --cacert=/etc/etcd/ca.pem --cert=/etc/etcd/kubernetes.pem --key=/etc/etcd/kubernetes-key.pem endpoint health --endpoints=https://192.168.1.21:2379,https://192.168.1.22:2379,https://192.168.1.23:2379 - Test VXLAN port accessibility between nodes (e.g., from worker1 to master1):
nc -zv 192.168.1.21 8472
If any of these tests fail, check your firewall rules or security groups to ensure the ports are open.
4. Verify Flannel and CoreDNS Status
After applying the fixes, check the pod status:
kubectl get pods -n kube-system
You should see Flannel pods enter the Running state, followed by CoreDNS pods transitioning to Running once the pod network is fully operational.
内容的提问来源于stack exchange,提问作者Mahdi Khosravi

