修复EmptyDir容量超限后K8s Pod错误提示未消失的解决问询
Hey there, let's tackle that stubborn error message hanging around even after you've fixed the storage usage issue. Here's what you can do to clear it up:
Step 1: Verify the Current Storage Usage First
Before making any changes, double-check that the agent EmptyDir volume is actually under the 100Mi limit now. Run these commands to confirm usage from within the pod:
# Check usage in the main container's mounted path kubectl exec -n dabai-fat soa-task-745d48d955-bd4j8 -- du -sh /opt/skywalking/agent # Or check the init container's path (they share the same volume) kubectl exec -n dabai-fat soa-task-745d48d955-bd4j8 -c init-agent -- du -sh /skywalking/agent
If the output shows usage below 100Mi, we can move on to refreshing the pod's state.
Step 2: Restart the Pod (Most Reliable Fix)
EmptyDir volumes are tied directly to the pod's lifecycle, so restarting the pod will create a fresh EmptyDir and clear any cached error states from Kubelet. You have two safe options:
- Rollout restart the deployment (maintains availability, ideal for production):
kubectl rollout restart deployment -n dabai-fat soa-task - Delete the specific pod (the ReplicaSet will automatically spin up a replacement):
kubectl delete pod -n dabai-fat soa-task-745d48d955-bd4j8
Once the new pod starts, run kubectl describe pod <new-pod-name> -n dabai-fat — the old error should no longer appear.
Step 3: Refresh Kubelet State (If You Can't Restart the Pod)
If restarting the pod isn't feasible right now, you can try restarting the Kubelet service on the node where the pod is running (azshara-k8s03). This forces Kubelet to recheck volume usage and update the pod's status:
- SSH into the node:
ssh azshara-k8s03 - Restart Kubelet:
sudo systemctl restart kubelet
⚠️ Heads up: Restarting Kubelet will temporarily disrupt all pods on the node, so only do this if the node has redundant workloads or you can tolerate a brief outage.
Step 4: Check External Monitoring Tools
If the error is showing up in an external monitoring/alerting tool (like Prometheus, Grafana, or your cluster's built-in alert system), the issue might be cached data. Wait for the tool's data refresh interval to pass, or manually trigger a refresh if the tool supports it.
内容的提问来源于stack exchange,提问作者Dolphin

