Kubernetes多Master节点Canal CNI启动失败问题求助
Let's break down your problem and walk through the solutions step by step, since you've got a solid background with K8s deployments but hit a snag with proxy and Canal on your multi-master cluster.
Root Cause Analysis
The core issue here is that your Master2 and Master3 nodes aren't getting assigned a Pod CIDR, which Flannel (part of Canal) requires to acquire a network lease. This usually stems from one of three key issues:
- kube-controller-manager misconfiguration: The control plane component responsible for assigning Pod CIDRs either has an incorrect
cluster-cidrvalue, or isn't enabled to allocate CIDRs to nodes. - Kubeadm configuration not taking effect: Even after editing the
kubeadm-configConfigMap, critical control plane components (apiserver, controller-manager) might not have picked up the updated settings. - Proxy interference: Your NO_PROXY setup could be missing critical internal cluster addresses, preventing the controller-manager from communicating with nodes or etcd to complete CIDR allocation.
Step-by-Step Solutions
1. Verify kube-controller-manager Configuration
First, ensure the controller-manager is properly configured to handle CIDR allocation:
- Check the static pod manifest at
/etc/kubernetes/manifests/kube-controller-manager.yamlfor these mandatory flags:command: - kube-controller-manager - --allocate-node-cidrs=true # Must be enabled to assign CIDRs to nodes - --cluster-cidr=192.168.151.0/25 # Exact match to your configured cluster CIDR # ... other existing flags - Confirm the NO_PROXY environment variable is set correctly inside the manifest (to avoid proxy blocking internal cluster traffic):
env: - name: NO_PROXY value: "127.0.0.1,localhost,192.168.151.0/25,<all-node-internal-ips>,<service-cidr>" - After modifying the manifest, delete the existing controller-manager pod to force a restart (kubelet will automatically recreate it with the new config):
kubectl delete pod -n kube-system -l component=kube-controller-manager - Check the controller-manager logs to confirm no connection or proxy-related errors:
kubectl logs -n kube-system -l component=kube-controller-manager
2. Manually Assign Pod CIDRs to Nodes
If automatic allocation isn't working immediately, you can manually split your /25 CIDR block and assign subnets to your master nodes:
- First, confirm the nodes have no Pod CIDR assigned:
kubectl get nodes -o wide - Assign a subnet to Master2 (e.g.,
192.168.151.0/26):kubectl patch node devmn2.cpdprd.pt -p '{"spec":{"podCIDR":"192.168.151.0/26"}}' - Assign the remaining subnet to Master3 (e.g.,
192.168.151.64/26):kubectl patch node devmn3.cpdprd.pt -p '{"spec":{"podCIDR":"192.168.151.64/26"}}' - Restart the Canal pods to pick up the new CIDRs:
kubectl delete pod -n kube-system -l k8s-app=canal
3. Fix Kubeadm Configuration Inconsistencies
Your note about kubeadm config view missing settings (like etcd cert paths) suggests the control plane isn't using your updated kubeadm-config ConfigMap:
- First, verify the ConfigMap has your correct settings:
kubectl get configmap kubeadm-config -n kube-system -o yaml - Ensure the
ClusterConfigurationsection includes your etcd and networking config:apiVersion: kubeadm.k8s.io/v1beta3 kind: ClusterConfiguration clusterName: <your-cluster-name> etcd: external: endpoints: ["<your-etcd-endpoint>"] caFile: <path-to-etcd-ca-cert> certFile: <path-to-etcd-client-cert> keyFile: <path-to-etcd-client-key> networking: podSubnet: 192.168.151.0/25 - To apply the ConfigMap changes, restart the apiserver and controller-manager pods (they're static pods, so deleting them triggers a rebuild with the new config):
kubectl delete pod -n kube-system -l component=kube-apiserver kubectl delete pod -n kube-system -l component=kube-controller-manager - Recheck
kubeadm config viewto confirm your settings are now visible.
4. Validate Proxy Settings
Double-check your NO_PROXY environment variable on all nodes (kubelet, control plane components) to ensure it includes:
- All node internal IPs
- The Pod CIDR (
192.168.151.0/25) - The Service CIDR (default is
10.96.0.0/12, or your custom range) - Localhost (
127.0.0.1) - Cluster domain (
cluster.local)
A missing entry here can block control plane components from communicating with nodes or etcd, preventing successful CIDR allocation.
Final Checks
After completing these steps:
- Verify nodes have Pod CIDRs assigned with
kubectl get nodes -o wide - Check Canal pod status with
kubectl get pods -n kube-system -l k8s-app=canal - Review kube-flannel logs again to confirm no lease acquisition errors:
kubectl logs -n kube-system -l k8s-app=canal -c kube-flannel
内容的提问来源于stack exchange,提问作者Zed_Blade

