如何检查GKE集群中Pod引导磁盘的空间占用情况?
Got it, let's figure out what's eating up your boot disk space on GKE. Here's a step-by-step approach to dig into the details:
1. 进入目标节点的调试容器
First, you can use kubectl debug to access the node showing disk pressure—no extra SSH setup needed:
kubectl debug node/<your-node-name> -it --image=ubuntu:latest
Replace <your-node-name> with the name of the node triggering alerts (get the full list with kubectl get nodes). Once inside the debug container, the node's root filesystem is mounted, so you can inspect disk usage directly.
2. Confirm the boot disk mount point
GKE's boot disk is typically mounted at /, but verify this with a quick disk check:
df -h
Look for the line with mount point /—this is your boot disk, and you'll see its total capacity and used space here to confirm it's indeed the source of the pressure.
3. Quickly spot large directories
Use the du command to scan top-level directories and identify which ones are hogging space:
du -h --max-depth=1 /
Focus on these common high-usage directories first—they're almost always the culprits:
/var/lib/containerdor/var/lib/docker: Stores container images, layers, logs, and temporary data (the biggest disk hog 9 times out of 10)/var/log: System and component logs (like kubelet, containerd) can balloon rapidly if log rotation isn't configured properly/var/lib/kubelet: Holds Pod-related files, including leftover data from terminated Pods that weren't cleaned up automatically
4. Drill down into specific directories
If, say, /var/lib/containerd is using too much space, dig deeper to find the exact source:
du -h --max-depth=1 /var/lib/containerd
- Check cached images: Use
crictl images(GKE uses containerd by default, so this is more reliable thandocker images) to see which large images are taking up space - Find unused containers: Run
crictl ps -ato spot terminated containers that are still lingering—their layers and logs can eat up disk space over time - Check log sizes: For Pod logs, run
du -h /var/log/pods/*to see if any specific Pod's logs have grown excessively large
5. Check GKE-specific system components
Don't overlook system-level logs and caches that might be hidden:
- Check systemd journal size:
journalctl --disk-usagewill show how much space system logs are using. If it's too big, you can configure log rotation or clear old logs safely. - Clean up unused images: Use
crictl rmi --pruneto remove unused images automatically (this won't delete images used by running Pods, so it's safe to run)
6. Quick overview without entering the node
If you just want a high-level view of all nodes' disk usage first, run:
kubectl top node | awk '{print $1, $4, $5}'
This gives you a quick glance at each node's disk usage percentage and available space, but you'll still need to enter the node for detailed directory-level checks to fix the issue.
内容的提问来源于stack exchange,提问作者Javier Guzmán

