AWS EKS扩缩容Pod时出现ALB 502错误该如何排查优化?
EKS ALB 峰值502错误排查与优化方案
排查步骤
- 匹配时间窗对应日志:将ALB访问日志、Pod生命周期事件、集群扩缩容事件的时间点对齐,确认502是出现在新Pod扩容上线阶段还是旧Pod缩容终止阶段
- 检查健康检查路径一致性:你当前配置的Pod readiness probe路径为
/_healthcheck/,但ALB健康检查路径为/healthcheck/,需先确认是否为配置手误,两端路径必须完全匹配 - 验证ALB目标组同步状态:峰值时aws-load-balancer-controller可能存在同步延迟,可检查对应时间点目标组内的Pod IP是否与当前就绪的Pod IP一致,是否存在已终止Pod未及时注销、未就绪Pod提前注册的情况
- 查看Pod终止阶段的流量处理逻辑:确认是否配置了preStop钩子与合理的终止宽限期,多数缩容阶段的502都是Pod被直接终止时ALB仍在向其转发请求导致
核心优化配置
1. 完善Pod生命周期配置
补充探针参数、preStop钩子与终止宽限期,示例配置如下:
# Pod模板配置 readinessProbe: httpGet: path: /_healthcheck/ port: 80 initialDelaySeconds: 5 # 适配应用启动耗时调整 periodSeconds: 2 timeoutSeconds: 1 successThreshold: 2 failureThreshold: 3 lifecycle: preStop: exec: command: ["sleep", "30"] # 时长需大于ALB健康检查间隔+注销延迟总和 terminationGracePeriodSeconds: 60 # 必须大于preStop的sleep时长
2. 调整ALB Ingress注解参数
优化健康检查逻辑,开启连接耗尽配置:
alb.ingress.kubernetes.io/healthcheck-path: "/_healthcheck/" # 和Pod探针路径保持一致 alb.ingress.kubernetes.io/healthcheck-interval-seconds: "2" alb.ingress.kubernetes.io/healthcheck-timeout-seconds: "1" alb.ingress.kubernetes.io/healthy-threshold-count: "2" alb.ingress.kubernetes.io/unhealthy-threshold-count: "2" # 开启目标注销时的连接耗尽,等待存量请求处理完成 alb.ingress.kubernetes.io/target-group-attributes: deregistration_delay.timeout_seconds=30
3. 峰值场景适配优化
如果集群规模大、扩缩容频率高,可调高aws-load-balancer-controller的并发同步worker数量,降低Pod状态到ALB目标组的同步延迟。
内容的提问来源于stack exchange,提问作者Most Wanted
相关产品推荐
相关产品推荐

