如何修改Kubernetes特定节点的PodCIDR?Calico CNI新节点报错求助
Let’s break down your problem and fix it step by step—no need to re-run kubeadm init, so you won’t risk breaking your existing working cluster. The core issue here is a mismatch between your Calico CNI’s configured pod network range, the pod CIDR assigned to new worker nodes, and kubelet’s settings.
1. First: Confirm Calico’s Pod Network Configuration
Before making any changes, verify what pod network range Calico is using. This is the baseline we need all nodes to align with:
- Check Calico’s ConfigMap:
Look for thekubectl get configmap calico-config -n kube-system -o yamlcalico_ipv4pool_cidrfield—this is the root range Calico uses to assign pod IPs. - Or check Calico’s IP Pool custom resource:
Note thekubectl get ippools.crd.projectcalico.org -o yamlcidrvalue in the output.
2. Modify PodCIDR for Specific Nodes
If a node has an incorrect PodCIDR (causing non-expected pod IPs and CNI errors), follow these steps to fix it:
Step 1: Cordon and Drain the Node
First, prevent new pods from scheduling to the node and evict existing pods (ensure your pods are managed by controllers like Deployments so they’ll restart elsewhere):
kubectl cordon <your-node-name> kubectl drain <your-node-name> --ignore-daemonsets
Step 2: Edit the Node’s PodCIDR
Update the node’s assigned pod CIDR to match a subnet within Calico’s root range:
kubectl edit node <your-node-name>
Find the spec.podCIDR field, replace its value with a valid subnet from Calico’s range (e.g., if Calico uses 192.168.0.0/16, use something like 192.168.10.0/24—make sure this subnet isn’t used by another node).
Step 3: Restart Kubelet on the Node
Apply the change by restarting kubelet:
systemctl restart kubelet
Step 4: Uncordon the Node
Allow pods to schedule back to the node once it’s ready:
kubectl uncordon <your-node-name>
3. Fix the "NetworkPlugin cni failed to set up pod" Error
If the CNI error persists after updating the PodCIDR, check these additional points:
- Verify the Calico DaemonSet is running on the node:
Ensure the pod for your target node is inkubectl get pods -n kube-system -o wide | grep calico-nodeRunningstate. - Check Calico node logs for configuration conflicts:
Look for errors related to IP range conflicts or missing network rules.kubectl logs <calico-node-pod-name> -n kube-system - Confirm node-to-node connectivity: Ensure the new worker node can communicate with the control plane and other nodes over Calico’s required ports (e.g., 4789 for IPIP, 179 for BGP).
4. Prevent This Issue for Future Worker Nodes
To avoid repeating this problem when adding new nodes, update your cluster’s controller manager to assign correct PodCIDRs automatically:
- Edit the kube-controller-manager static pod manifest (usually located at
/etc/kubernetes/manifests/kube-controller-manager.yamlon control plane nodes):vi /etc/kubernetes/manifests/kube-controller-manager.yaml - Add or modify the
--cluster-cidrflag in the container command section to match Calico’s root pod range:spec: containers: - command: - kube-controller-manager - --cluster-cidr=192.168.0.0/16 # Replace with your Calico CIDR # Keep all other existing command flags - Save the file—kubelet will automatically restart the kube-controller-manager pod with the new setting. Future nodes joined via
kubeadm joinwill now get valid PodCIDRs aligned with Calico.
内容的提问来源于stack exchange,提问作者cryptoparty

