求EKS集群单副本Pod节点自动扩缩容时的零停机解决方案
解决EKS中Cluster Autoscaler单副本Pod迁移时的服务中断问题
针对EKS集群中Cluster Autoscaler节点扩缩容导致单副本Pod迁移时服务中断(旧Pod提前终止、新Pod未就绪)的问题,可通过以下关键配置组合解决:
1. 配置Pod Disruption Budget(PDB)约束自愿中断行为
Cluster Autoscaler的节点驱逐属于自愿中断,PDB能确保集群在这类操作中维持指定数量的可用Pod。单副本场景下设置minAvailable: 1,强制Cluster Autoscaler必须等新Pod完全就绪后,才能终止旧Pod。
示例PDB配置:
apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: app-pdb spec: minAvailable: 1 selector: matchLabels: app: your-app # 替换为你的Pod标签
2. 调整Deployment滚动更新策略
将Deployment的滚动更新策略设置为maxUnavailable: 0,确保Pod重建(含节点驱逐触发的重建)过程中始终有一个可用Pod运行,旧Pod不会在新Pod就绪前被终止。
示例Deployment策略配置:
apiVersion: apps/v1 kind: Deployment metadata: name: your-app-deployment spec: replicas: 1 strategy: rollingUpdate: maxSurge: 1 # 允许额外创建1个Pod maxUnavailable: 0 # 不允许任何Pod不可用 type: RollingUpdate selector: matchLabels: app: your-app template: metadata: labels: app: your-app spec: containers: - name: app-container image: your-app-image:latest ports: - containerPort: 80 readinessProbe: # 必须配置就绪探针,确保K8s能准确判断Pod是否就绪 httpGet: path: /healthz port: 80 initialDelaySeconds: 10 periodSeconds: 5
3. 配置Pod生命周期钩子(可选,额外保障)
添加preStop钩子延迟旧Pod的终止时间,给负载均衡器足够时间将流量切换到新Pod,避免短暂502响应。同时配合terminationGracePeriodSeconds确保宽限期足够。
示例生命周期钩子配置:
spec: containers: - name: app-container # ... 其他配置 ... lifecycle: preStop: exec: command: ["sleep", "20"] # 根据实际流量切换时间调整 terminationGracePeriodSeconds: 40 # 需大于preStop延迟时间,确保钩子执行完成
注意事项
- 必须为Pod配置就绪探针:K8s依赖就绪探针判断Pod是否真正可用,无就绪探针时,即使容器启动,K8s也无法确认服务就绪,会导致旧Pod提前终止。
- 确保Cluster Autoscaler版本≥1.14:旧版本对PDB的支持不完善,建议使用EKS官方推荐的稳定版本。
- 验证配置:部署完成后,可手动标记节点为不可调度并驱逐Pod,观察旧Pod是否在新Pod进入
Ready状态后才切换为Terminating。
内容的提问来源于stack exchange,提问作者Kelvin Wong
相关产品推荐
相关产品推荐

