You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置Kubernetes自动修复部署?缩短节点故障自愈时长

Shortening Pod Rescheduling Time After Node Failure in Kubernetes 1.10

Hey there! That 5-minute wait for pods to reschedule after a node goes down is way too slow for most production scenarios. Let's break down the key Kubernetes components and parameters you can tweak to cut this down significantly:

1. Adjust Kube-Controller-Manager Parameters

The main culprit behind that 5-minute delay is the default pod eviction timeout. Here are the two critical flags to modify:

  • --pod-eviction-timeout: This is how long the controller manager waits after marking a node as NotReady before evicting its pods. The default is 5m0s—this is exactly the delay you're seeing. You can lower this to something like 30s or 1m0s depending on your tolerance for false positives (like temporary network blips).
  • --node-monitor-grace-period: This defines how long the controller manager goes without receiving node heartbeats before marking the node as NotReady. The default is 40s; you can reduce this to 20s to speed up failure detection, but don't go too low (under 10s) to avoid flagging healthy nodes during transient network issues.

To apply these changes, edit the kube-controller-manager's manifest (usually in /etc/kubernetes/manifests/ on control plane nodes) or adjust its startup arguments, then restart the component.

2. Tune Kubelet Status Reporting

The kubelet sends node status updates to the API server periodically. Tweaking this frequency helps the controller detect failures faster:

  • --node-status-update-frequency: This controls how often the kubelet reports its status to the API server. The default is 10s. You can lower it to 5s to reduce the time it takes for the API server to notice a node has stopped responding. Just balance this with the extra load it might add to your API server.

3. Optional: Speed Up Pod Termination

While not directly related to rescheduling delay, reducing terminationGracePeriodSeconds in your pod specs can make the overall failover feel quicker. This parameter defines how long Kubernetes waits for a pod to shut down gracefully before force-killing it. For stateless apps, you could set this to 10s or even 5s (just make sure your app handles quick shutdowns properly).

apiVersion: apps/v1
kind: Deployment
metadata:
  name: your-app
spec:
  template:
    spec:
      terminationGracePeriodSeconds: 10
      containers:
      - name: your-container
        image: your-image:tag

Important Notes

  • Test first: Always tweak these parameters in a staging environment before applying to production. Too-aggressive timeouts can lead to unnecessary pod evictions during network flutters.
  • Kubernetes 1.10 specifics: These flags are valid for 1.10, and while newer versions have evolved some failure detection logic, the core principles here still hold.

内容的提问来源于stack exchange,提问作者Nurza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:14:43