AWS EKS集群Pod ReadinessGate适配Target Group失效排查:报错'target-health.alb.ingress.k8s.aws条件不存在'
Let's break down your issues step by step, focusing on compatibility, configuration gaps, and actionable fixes:
1. EKS 1.20 vs ALB Ingress Controller 1.1.4: Compatibility Breakage
First off, EKS 1.20 is too new for ALB Ingress Controller 1.1.4—this is likely the root cause of your ReadinessGate failure. The 1.1.x series was built for older Kubernetes versions (roughly 1.16-1.19) and doesn't support the API changes and Pod condition handling introduced in 1.20.
When you see the error corresponding condition of pod readiness gate "target-health.alb.ingress.k8s.aws/afik-nginx-ingress_afik-nginx-service_80" does not exist, it means the controller can't recognize or update the required Pod condition because it's incompatible with EKS 1.20's underlying mechanisms.
2. Configuration Gaps in Your 1.1.x Setup
Even if version compatibility wasn't an issue, your current setup has a few potential pitfalls:
- Cross-namespace resource monitoring: Your ALB Controller runs in
kube-system, but your test resources are inafik-test. By default, the 1.1.x controller only watches the namespace it's deployed in unless you explicitly set the--watch-namespaceflag to includeafik-test. Without this, the controller won't process your Ingress/Service/Deployment, so it never adds the required readiness condition to your Pods. - Headless Service limitation: You're using a
ClusterIP: NoneHeadless Service. The 1.1.x controller has known bugs with ReadinessGate support for headless services—switch to a standard ClusterIP Service for testing. - Deprecated Ingress API: Your Ingress uses
extensions/v1beta1, which was deprecated in Kubernetes 1.19 and removed in 1.22. While EKS 1.20 still allows it, the 1.1.x controller may struggle to parse this API version correctly, leading to failed condition updates.
3. Workarounds for Zero-Downtime Deployments (With 1.1.x)
If you can't upgrade immediately, try these temporary fixes to mitigate 502/504 errors:
- Extend termination grace period: Add
terminationGracePeriodSeconds: 120to your Deployment spec. This gives the controller more time to deregister old Pods from the Target Group before they're terminated. - Add a preStop hook: Delay Pod termination to let the controller sync with AWS:
containers: - name: nginx image: nginx ports: - containerPort: 80 lifecycle: preStop: exec: command: ["sleep", "30"] - Reduce controller sync interval: Modify your ALB Controller Deployment to set
--sync-period=60s(down from the default 300s). This makes the controller check for Target Group updates more frequently, cutting down on registration/deregistration delays.
4. Long-Term Fix: Upgrade to ALB Ingress Controller 2.x
This is the most reliable solution to resolve all your issues:
- 2.x versions are fully compatible with EKS 1.20+ and handle Pod readiness conditions, Target Group sync, and cross-namespace resources seamlessly.
- It includes native optimizations for zero-downtime deployments, with simpler ReadinessGate configuration (or even auto-handling in some cases).
- When upgrading:
- Uninstall the 1.1.4 controller completely.
- Install the latest 2.x version, following standard setup procedures for your cluster.
- Update your Ingress resources to use
networking.k8s.io/v1instead ofextensions/v1beta1, and adjust any annotations that changed between versions.
内容的提问来源于stack exchange,提问作者Afik.A

