You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何排查GKE中抢占式节点(preemptible nodes)频繁重置问题?

Troubleshooting Unexpected Preemptible Node Terminations in GKE Dev Cluster

Let's break down your problem and walk through actionable steps to figure out why you're seeing 3-4 daily health check alerts instead of the expected 1 per node's 24-hour lifecycle.

Possible Root Causes

First, let's rule out the most likely culprits for frequent node replacements:

  • Unplanned Preemption: GCP preemptible nodes aren't just terminated after 24 hours—they can also be reclaimed early if GCP needs capacity for standard instances. Even low-spec f1-micro nodes aren't immune to this, especially in busy regions like us-west1.
  • Node Autorepair Triggers: Your f1-micro nodes have extremely limited resources (0.2vCPU, 0.6GB RAM). If pods consume too much resources, kubelet might become unresponsive, or the node might fail GKE's health checks. Autorepair will automatically replace unhealthy nodes, which would trigger your alerts.
  • Autoscaling Fluctuations: If your workload's pod count spikes and drops frequently, the node pool might scale up then scale down. While scaling down usually targets idle nodes, if a node has pods that can't be evicted quickly, it might cause temporary health check failures.
  • Node Autoupgrades: While less likely (autoupgrades typically run once per minor version), if there are frequent patch updates, you might see multiple node replacements in a day.

How to Find the Smoking Gun in Logs

You're right that GKE/Cloud Logging can feel overwhelming—here's how to narrow down the relevant logs:

1. Cloud Logging (Formerly Stackdriver) Queries

Use these targeted queries in the Cloud Logging Console to find node termination events:

  • Node Termination & Preemption Logs:
    resource.type="k8s_node"
    resource.labels.cluster_name="dev"
    resource.labels.node_pool_name="default-pool-f1-micro-preemptible"
    (textPayload:"preempted" OR textPayload:"autorepair" OR textPayload:"upgrade" OR textPayload:"terminated" OR textPayload:"shutdown")
    
  • Compute Engine VM Instance Logs:
    resource.type="gce_instance"
    resource.labels.instance_name=~"gke-dev-default-pool-f1-micro-preemptible.*"
    (textPayload:"preempt" OR textPayload:"terminate" OR textPayload:"autorepair")
    
  • Health Check Failure Logs:
    resource.type="http_load_balancer"
    textPayload:"health check failed"
    
    Match the failed pod/node IPs from these logs to the node termination timestamps.

2. Kubernetes Cluster Events

Run this command to get a chronological list of cluster events—look for NodePreemption, NodeTermination, or SuccessfulDelete events tied to your preemptible nodes:

kubectl get events --all-namespaces --sort-by='.metadata.creationTimestamp'

This will show you exactly when nodes were deleted, and why (e.g., "Node is marked for deletion due to autorepair").

3. Node OS Logs (If Nodes Are Still Running)

If you catch a node before it's terminated, exec into it to check system logs for shutdown triggers:

kubectl exec -it <node-name> -- cat /var/log/syslog | grep -i "preempt\|shutdown\|kubelet"

This can show if kubelet crashed, or if the node received a preemption signal from GCP.

Actionable Troubleshooting Steps

  1. Monitor Node Resource Usage in Real Time
    f1-micro nodes are easy to overload. Run kubectl top nodes to check CPU/memory utilization. If nodes are consistently hitting 90%+ usage, your pods are likely causing resource starvation, which triggers autorepair. Fix this by adding resource requests and limits to your pods to prevent them from hogging resources.

  2. Verify Autorepair Triggers
    Check your node pool's autorepair configuration with:

    gcloud container node-pools describe default-pool-f1-micro-preemptible --cluster=dev --zone=us-west1
    

    Look for autorepair settings. You can also temporarily disable autorepair (for a few hours) to see if alert frequency drops—if it does, you know autorepair is the culprit.

  3. Track Node Lifespans
    In the Compute Engine Console, filter for your preemptible nodes and check their "Creation time" vs "Deletion time". If nodes are being deleted well before the 24-hour mark, that's a sign of unplanned preemption.

  4. Test With a Higher-Spec Node (Temporarily)
    To rule out resource constraints, create a small temporary node pool with g1-small nodes (still preemptible) and move a few workloads there. If alerts stop, you know the f1-micro nodes are under-resourced for your workload.

Final Notes

You don't need to destroy your cluster to debug this—focus on isolating the cause with logs and targeted tests first. Once you identify the root issue (preemption vs autorepair vs resource limits), you can adjust your configuration accordingly (e.g., add resource limits, switch to a slightly larger machine type, or adjust autorepair thresholds).

内容的提问来源于stack exchange,提问作者Paul Reimer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:51:37