You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生产环境Kubernetes节点NotReady/NodeLost问题求助:磁盘IO占用过高

Kubernetes Agent Nodes Entering NotReady State (Pod NodeLost) Due to High Disk Resource Usage - Troubleshooting & Recommendations

Hey Hugo, sorry to hear you're stuck with this critical production cluster issue—let's walk through actionable steps to diagnose the root cause and find a fix, beyond just cranking up IOPS.

1. First: Pinpoint Exactly What's Eating Disk Space

High disk usage is a symptom, not the root cause. Let's narrow down the culprit:

  • Log into an affected node (if still accessible) and run these commands to spot large directories/files:
    # Check overall disk usage across mounts
    df -h
    # Drill into core Kubernetes/container storage directories
    du -sh /var/lib/kubelet/* /var/lib/containerd/*  # Adjust path to /var/lib/docker/* if using Docker
    
  • Container logs are a top offender: Verify your container runtime's log rotation settings. For containerd, check /etc/containerd/config.toml for log size/retention limits; for Docker, look at /etc/logrotate.d/docker. Unrotated logs can fill disks rapidly even with moderate traffic.
  • Unused container images: Kubernetes doesn't auto-clean old images by default. Run crictl images (containerd) or docker images to list unused images, then trigger a manual cleanup with:
    kubelet image gc
    
    You can also adjust Kubelet's imageGCHighThresholdPercent (default 85%) and imageGCLowThresholdPercent (default 80%) to make automatic cleanup kick in earlier.

2. Verify If IOPS Is Actually the Bottleneck

You bumped IOPS to 2000 but saw no improvement—let's check if the disk is truly saturated or if other factors are at play:

  • Monitor real-time disk performance on a node with:
    iostat -x 5
    
    Keep an eye on %util (if it hits 100%, the disk is saturated even if IOPS are under your provisioned limit) and await (average I/O latency; high values mean the disk can't keep up with requests).
  • Check if nodes are hitting Kubernetes' DiskPressure eviction threshold: Run kubectl describe node <node-name> and look at the Conditions section. If DiskPressure is marked True, the Kubelet is intentionally marking the node NotReady to protect it. Tweaking evictionHard parameters in Kubelet's config is a band-aid, not a long-term fix.
  • Confirm disk caching settings: For Premium disks attached to your Standard_DS2_VX VMs, ensure you're using the ReadWrite caching mode (optimal for Kubelet/container storage) instead of ReadOnly or None—this can drastically improve real-world IO performance.

3. ACS-Engine Specific Configuration Checks

Since your cluster is deployed with acs-engine 0.17.0, let's rule out misconfigurations in the cluster template:

  • Check if Kubelet/container runtime data is on the Premium disk: Verify /etc/fstab to ensure your Premium disk is mounted to /var/lib/kubelet and /var/lib/containerd (or Docker's root dir). If these directories are still on the small OS disk, upgrading the data disk won't help.
  • Review acs-engine's generated Kubelet config: On nodes, look at /var/lib/kubelet/config.yaml for parameters like evictionHard, volumeStatsAggPeriod, and imageGC settings—ensure they align with your cluster's storage needs.

4. Next Steps Before Trying Higher IOPS

Before you provision higher IOPS, try these to validate and mitigate:

  • Temporary recovery: If nodes are accessible, manually delete old logs, unused images, and terminated Pods to free up disk space. If nodes return to Ready state, you've confirmed disk pressure is the trigger.
  • Enforce Pod resource limits: Ensure all Pods define resources.requests.ephemeral-storage and resources.limits.ephemeral-storage to prevent rogue Pods from consuming all node disk space.
  • Correlate events with monitoring: Use your cluster's monitoring tooling (including the attached screenshots) to map disk usage spikes to specific events—e.g., a large deployment, image pull, or log burst. This will help you address the source of the high disk/IO load, not just the symptom.

内容的提问来源于stack exchange,提问作者Hugo Marcelo Del Negro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:38:07