You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes多Master节点Canal CNI启动失败问题求助

Troubleshooting Canal CrashLoopBackOff & Missing Pod CIDR on Master Nodes

Let's break down your problem and walk through the solutions step by step, since you've got a solid background with K8s deployments but hit a snag with proxy and Canal on your multi-master cluster.

Root Cause Analysis

The core issue here is that your Master2 and Master3 nodes aren't getting assigned a Pod CIDR, which Flannel (part of Canal) requires to acquire a network lease. This usually stems from one of three key issues:

  • kube-controller-manager misconfiguration: The control plane component responsible for assigning Pod CIDRs either has an incorrect cluster-cidr value, or isn't enabled to allocate CIDRs to nodes.
  • Kubeadm configuration not taking effect: Even after editing the kubeadm-config ConfigMap, critical control plane components (apiserver, controller-manager) might not have picked up the updated settings.
  • Proxy interference: Your NO_PROXY setup could be missing critical internal cluster addresses, preventing the controller-manager from communicating with nodes or etcd to complete CIDR allocation.

Step-by-Step Solutions

1. Verify kube-controller-manager Configuration

First, ensure the controller-manager is properly configured to handle CIDR allocation:

  • Check the static pod manifest at /etc/kubernetes/manifests/kube-controller-manager.yaml for these mandatory flags:
    command:
      - kube-controller-manager
      - --allocate-node-cidrs=true  # Must be enabled to assign CIDRs to nodes
      - --cluster-cidr=192.168.151.0/25  # Exact match to your configured cluster CIDR
      # ... other existing flags
    
  • Confirm the NO_PROXY environment variable is set correctly inside the manifest (to avoid proxy blocking internal cluster traffic):
    env:
      - name: NO_PROXY
        value: "127.0.0.1,localhost,192.168.151.0/25,<all-node-internal-ips>,<service-cidr>"
    
  • After modifying the manifest, delete the existing controller-manager pod to force a restart (kubelet will automatically recreate it with the new config):
    kubectl delete pod -n kube-system -l component=kube-controller-manager
    
  • Check the controller-manager logs to confirm no connection or proxy-related errors:
    kubectl logs -n kube-system -l component=kube-controller-manager
    

2. Manually Assign Pod CIDRs to Nodes

If automatic allocation isn't working immediately, you can manually split your /25 CIDR block and assign subnets to your master nodes:

  • First, confirm the nodes have no Pod CIDR assigned:
    kubectl get nodes -o wide
    
  • Assign a subnet to Master2 (e.g., 192.168.151.0/26):
    kubectl patch node devmn2.cpdprd.pt -p '{"spec":{"podCIDR":"192.168.151.0/26"}}'
    
  • Assign the remaining subnet to Master3 (e.g., 192.168.151.64/26):
    kubectl patch node devmn3.cpdprd.pt -p '{"spec":{"podCIDR":"192.168.151.64/26"}}'
    
  • Restart the Canal pods to pick up the new CIDRs:
    kubectl delete pod -n kube-system -l k8s-app=canal
    

3. Fix Kubeadm Configuration Inconsistencies

Your note about kubeadm config view missing settings (like etcd cert paths) suggests the control plane isn't using your updated kubeadm-config ConfigMap:

  • First, verify the ConfigMap has your correct settings:
    kubectl get configmap kubeadm-config -n kube-system -o yaml
    
  • Ensure the ClusterConfiguration section includes your etcd and networking config:
    apiVersion: kubeadm.k8s.io/v1beta3
    kind: ClusterConfiguration
    clusterName: <your-cluster-name>
    etcd:
      external:
        endpoints: ["<your-etcd-endpoint>"]
        caFile: <path-to-etcd-ca-cert>
        certFile: <path-to-etcd-client-cert>
        keyFile: <path-to-etcd-client-key>
    networking:
      podSubnet: 192.168.151.0/25
    
  • To apply the ConfigMap changes, restart the apiserver and controller-manager pods (they're static pods, so deleting them triggers a rebuild with the new config):
    kubectl delete pod -n kube-system -l component=kube-apiserver
    kubectl delete pod -n kube-system -l component=kube-controller-manager
    
  • Recheck kubeadm config view to confirm your settings are now visible.

4. Validate Proxy Settings

Double-check your NO_PROXY environment variable on all nodes (kubelet, control plane components) to ensure it includes:

  • All node internal IPs
  • The Pod CIDR (192.168.151.0/25)
  • The Service CIDR (default is 10.96.0.0/12, or your custom range)
  • Localhost (127.0.0.1)
  • Cluster domain (cluster.local)
    A missing entry here can block control plane components from communicating with nodes or etcd, preventing successful CIDR allocation.

Final Checks

After completing these steps:

  1. Verify nodes have Pod CIDRs assigned with kubectl get nodes -o wide
  2. Check Canal pod status with kubectl get pods -n kube-system -l k8s-app=canal
  3. Review kube-flannel logs again to confirm no lease acquisition errors:
    kubectl logs -n kube-system -l k8s-app=canal -c kube-flannel
    

内容的提问来源于stack exchange,提问作者Zed_Blade

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:37:39