生产环境Kubernetes节点NotReady/NodeLost问题求助:磁盘IO占用过高
Kubernetes Agent Nodes Entering NotReady State (Pod NodeLost) Due to High Disk Resource Usage - Troubleshooting & Recommendations
Hey Hugo, sorry to hear you're stuck with this critical production cluster issue—let's walk through actionable steps to diagnose the root cause and find a fix, beyond just cranking up IOPS.
1. First: Pinpoint Exactly What's Eating Disk Space
High disk usage is a symptom, not the root cause. Let's narrow down the culprit:
- Log into an affected node (if still accessible) and run these commands to spot large directories/files:
# Check overall disk usage across mounts df -h # Drill into core Kubernetes/container storage directories du -sh /var/lib/kubelet/* /var/lib/containerd/* # Adjust path to /var/lib/docker/* if using Docker - Container logs are a top offender: Verify your container runtime's log rotation settings. For containerd, check
/etc/containerd/config.tomlfor log size/retention limits; for Docker, look at/etc/logrotate.d/docker. Unrotated logs can fill disks rapidly even with moderate traffic. - Unused container images: Kubernetes doesn't auto-clean old images by default. Run
crictl images(containerd) ordocker imagesto list unused images, then trigger a manual cleanup with:
You can also adjust Kubelet'skubelet image gcimageGCHighThresholdPercent(default 85%) andimageGCLowThresholdPercent(default 80%) to make automatic cleanup kick in earlier.
2. Verify If IOPS Is Actually the Bottleneck
You bumped IOPS to 2000 but saw no improvement—let's check if the disk is truly saturated or if other factors are at play:
- Monitor real-time disk performance on a node with:
Keep an eye oniostat -x 5%util(if it hits 100%, the disk is saturated even if IOPS are under your provisioned limit) andawait(average I/O latency; high values mean the disk can't keep up with requests). - Check if nodes are hitting Kubernetes' DiskPressure eviction threshold: Run
kubectl describe node <node-name>and look at theConditionssection. IfDiskPressureis markedTrue, the Kubelet is intentionally marking the node NotReady to protect it. TweakingevictionHardparameters in Kubelet's config is a band-aid, not a long-term fix. - Confirm disk caching settings: For Premium disks attached to your Standard_DS2_VX VMs, ensure you're using the ReadWrite caching mode (optimal for Kubelet/container storage) instead of
ReadOnlyorNone—this can drastically improve real-world IO performance.
3. ACS-Engine Specific Configuration Checks
Since your cluster is deployed with acs-engine 0.17.0, let's rule out misconfigurations in the cluster template:
- Check if Kubelet/container runtime data is on the Premium disk: Verify
/etc/fstabto ensure your Premium disk is mounted to/var/lib/kubeletand/var/lib/containerd(or Docker's root dir). If these directories are still on the small OS disk, upgrading the data disk won't help. - Review acs-engine's generated Kubelet config: On nodes, look at
/var/lib/kubelet/config.yamlfor parameters likeevictionHard,volumeStatsAggPeriod, andimageGCsettings—ensure they align with your cluster's storage needs.
4. Next Steps Before Trying Higher IOPS
Before you provision higher IOPS, try these to validate and mitigate:
- Temporary recovery: If nodes are accessible, manually delete old logs, unused images, and terminated Pods to free up disk space. If nodes return to Ready state, you've confirmed disk pressure is the trigger.
- Enforce Pod resource limits: Ensure all Pods define
resources.requests.ephemeral-storageandresources.limits.ephemeral-storageto prevent rogue Pods from consuming all node disk space. - Correlate events with monitoring: Use your cluster's monitoring tooling (including the attached screenshots) to map disk usage spikes to specific events—e.g., a large deployment, image pull, or log burst. This will help you address the source of the high disk/IO load, not just the symptom.
内容的提问来源于stack exchange,提问作者Hugo Marcelo Del Negro
相关产品推荐
相关产品推荐

