You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多Master高可用Kubernetes集群Flannel组件CrashLoopBackOff故障求助

Fixing Flannel CrashLoopBackOff and CoreDNS ContainerCreating Issues in HA Kubernetes Cluster

First off, let's start by diagnosing the root causes behind your Flannel and CoreDNS issues—since CoreDNS is stuck in ContainerCreating, that's almost certainly because the pod network (Flannel) isn't functioning properly. Let's break down the steps to resolve this:

1. Verify Pod CIDR Alignment Between Kubeadm and Flannel

Flannel defaults to using the 10.244.0.0/16 Pod CIDR, but looking at your config.yaml, you didn't specify this in your ClusterConfiguration. Kubernetes needs this to assign pod IPs correctly.

  • Check if the kube-controller-manager has the correct --cluster-cidr flag:
    kubectl get pods kube-controller-manager-master1 -n kube-system -o yaml | grep cluster-cidr
    
  • If the flag is missing, edit the static pod manifest on each master node:
    sudo vi /etc/kubernetes/manifests/kube-controller-manager.yaml
    
    Add this line under the command section:
    - --cluster-cidr=10.244.0.0/16
    
    The kubelet will automatically restart the controller-manager pod once you save the file.

2. Fix Flannel's Etcd Configuration (Critical for External Etcd)

Your external etcd cluster has client certificate authentication enabled, but the default Flannel manifest doesn't account for this. Flannel needs valid certificates to communicate with your etcd cluster, and it's currently trying to use the default local etcd endpoint (http://127.0.0.1:2379) which won't work.

Step 2.1: Customize the Flannel Manifest

  1. Download the original Flannel YAML:
    wget https://raw.githubusercontent.com/coreos/flannel/master/Documentation/kube-flannel.yml
    
  2. Edit the kube-flannel-cfg ConfigMap:
    Find the net-conf.json section and update it to include your etcd endpoints and certificate paths:
    {
      "Network": "10.244.0.0/16",
      "Backend": {
        "Type": "vxlan"
      },
      "EtcdEndpoints": "https://192.168.1.21:2379,https://192.168.1.22:2379,https://192.168.1.23:2379",
      "EtcdCAFile": "/etc/etcd/ca.pem",
      "EtcdCertFile": "/etc/etcd/kubernetes.pem",
      "EtcdKeyFile": "/etc/etcd/kubernetes-key.pem"
    }
    
  3. Add a volume mount for etcd certificates in the Flannel DaemonSet:
    • Under the volumes section of the DaemonSet, add:
      - name: etcd-certs
        hostPath:
          path: /etc/etcd
          type: Directory
      
    • Under the containers[0].volumeMounts section, add:
      - name: etcd-certs
        mountPath: /etc/etcd
        readOnly: true
      

Step 2.2: Reapply the Modified Flannel Manifest

  1. Delete the existing Flannel resources:
    kubectl delete -f kube-flannel.yml
    
  2. Apply the customized manifest:
    kubectl apply -f kube-flannel.yml
    

3. Validate Network Connectivity Between Nodes

Flannel uses VXLAN (UDP port 8472) for pod-to-pod communication, and your nodes need access to the etcd ports (TCP 2379/2380).

  • Test connectivity to etcd from a worker node:
    ETCDCTL_API=3 etcdctl --cacert=/etc/etcd/ca.pem --cert=/etc/etcd/kubernetes.pem --key=/etc/etcd/kubernetes-key.pem endpoint health --endpoints=https://192.168.1.21:2379,https://192.168.1.22:2379,https://192.168.1.23:2379
    
  • Test VXLAN port accessibility between nodes (e.g., from worker1 to master1):
    nc -zv 192.168.1.21 8472
    

If any of these tests fail, check your firewall rules or security groups to ensure the ports are open.

4. Verify Flannel and CoreDNS Status

After applying the fixes, check the pod status:

kubectl get pods -n kube-system

You should see Flannel pods enter the Running state, followed by CoreDNS pods transitioning to Running once the pod network is fully operational.

内容的提问来源于stack exchange,提问作者Mahdi Khosravi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 14:37:47