请求协助排查Kubernetes集群不稳定及节点、Pod异常问题
Hey there, let's tackle each of these cluster problems one by one with practical steps to get things stabilized—especially since you're rolling out more services, resolving these quickly will prevent bigger headaches down the line:
1. System OOM Encountered on Nodes
System OOM means your node's memory is completely exhausted, so let's track down the root cause:
- Check node resource usage: Run
kubectl top nodesto see which nodes are maxing out on memory. Then usekubectl describe node <node-name>to look at theAllocated resourcessection andConditions(likeMemoryPressure) to confirm resource exhaustion. - Dig into OOM logs: On the affected node, run
dmesg | grep -i oomorjournalctl -u kubelet | grep -i oomto find which process was killed by the OOM killer. This will tell you if it's a system process or a misbehaving pod hogging memory. - Enforce pod resource limits: A common culprit is pods without
resources.limitsset, letting them consume unlimited memory. Use this command to spot pods missing memory limits:
Add reasonablekubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}/{.metadata.name}: {.spec.containers[*].resources.limits.memory}{"\n"}' | grep -E ": $|:\s*<none>"limitsandrequeststo all pods to prevent them from starving system resources. - Verify node resource reservations: Make sure your kubelet is configured with
--kube-reservedand--system-reservedflags to set aside memory for system and kubernetes processes—this ensures pods can't use all the node's memory.
2. Non-Production Cluster Pod Scheduling Imbalance
If all pods are landing on a single node, here's how to fix the spread:
- Check node taints and pod tolerations: Run
kubectl describe node <node-name>for each node to see if other nodes have taints that pods can't tolerate. Then check your pod specs fortolerations—maybe they're only set to tolerate the single node's taints. - Confirm node readiness: Run
kubectl get nodesto ensure all nodes are inReadystatus. If any showSchedulingDisabled, that's why pods aren't being scheduled there. - Audit pod affinity rules: Check if your pods have
nodeAffinityrules forcing them to a specific node, or misconfiguredpodAntiAffinitythat prevents scheduling on other nodes. - Check scheduler health: Look at the kube-scheduler logs with
kubectl logs -n kube-system kube-scheduler-<control-plane-node>to see if there are scheduling errors or misconfigurations causing the imbalance.
3. Pods Stuck in 'Waiting: ContainerCreating' State
This usually points to a problem with setup before the container starts—here's how to diagnose:
- Inspect pod events: The first step is always
kubectl describe pod <pod-name> -n <namespace>. Look at theEventssection—it will tell you exactly what's failing: image pull errors, volume mount issues, network setup problems, etc. - Validate image pull: If the event mentions
ImagePullBackOfforErrImagePull, confirm the image URL is correct, the node can reach the registry, and you have the right credentials (like a secret for private registries). - Check storage volumes: If you're using persistent volumes, verify the PVC is bound to a PV with
kubectl get pvc -n <namespace>. Ensure the storage backend (NFS, cloud disk, etc.) is accessible and healthy. - Troubleshoot network plugin: Make sure your CNI plugin (Calico, Flannel, etc.) is running properly. Check the CNI pods in the kube-system namespace with
kubectl get pods -n kube-system | grep cni—if any are crashing, that's likely the issue. Also verify the node's CNI configuration files are intact.
Since you're deploying more services, fixing these foundational issues will make your cluster more resilient as you scale. Start with the OOM problem first (it's the most critical for node stability), then work through the scheduling and stuck pod issues. If you hit specific error messages during troubleshooting, feel free to share them for more targeted help!
内容的提问来源于stack exchange,提问作者Darshan Patel

