You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS EKS集群Pod ReadinessGate适配Target Group失效排查:报错'target-health.alb.ingress.k8s.aws条件不存在'

问题分析与解决方案

Let's break down your issues step by step, focusing on compatibility, configuration gaps, and actionable fixes:

1. EKS 1.20 vs ALB Ingress Controller 1.1.4: Compatibility Breakage

First off, EKS 1.20 is too new for ALB Ingress Controller 1.1.4—this is likely the root cause of your ReadinessGate failure. The 1.1.x series was built for older Kubernetes versions (roughly 1.16-1.19) and doesn't support the API changes and Pod condition handling introduced in 1.20.

When you see the error corresponding condition of pod readiness gate "target-health.alb.ingress.k8s.aws/afik-nginx-ingress_afik-nginx-service_80" does not exist, it means the controller can't recognize or update the required Pod condition because it's incompatible with EKS 1.20's underlying mechanisms.

2. Configuration Gaps in Your 1.1.x Setup

Even if version compatibility wasn't an issue, your current setup has a few potential pitfalls:

  • Cross-namespace resource monitoring: Your ALB Controller runs in kube-system, but your test resources are in afik-test. By default, the 1.1.x controller only watches the namespace it's deployed in unless you explicitly set the --watch-namespace flag to include afik-test. Without this, the controller won't process your Ingress/Service/Deployment, so it never adds the required readiness condition to your Pods.
  • Headless Service limitation: You're using a ClusterIP: None Headless Service. The 1.1.x controller has known bugs with ReadinessGate support for headless services—switch to a standard ClusterIP Service for testing.
  • Deprecated Ingress API: Your Ingress uses extensions/v1beta1, which was deprecated in Kubernetes 1.19 and removed in 1.22. While EKS 1.20 still allows it, the 1.1.x controller may struggle to parse this API version correctly, leading to failed condition updates.

3. Workarounds for Zero-Downtime Deployments (With 1.1.x)

If you can't upgrade immediately, try these temporary fixes to mitigate 502/504 errors:

  • Extend termination grace period: Add terminationGracePeriodSeconds: 120 to your Deployment spec. This gives the controller more time to deregister old Pods from the Target Group before they're terminated.
  • Add a preStop hook: Delay Pod termination to let the controller sync with AWS:
    containers:
    - name: nginx
      image: nginx
      ports:
      - containerPort: 80
      lifecycle:
        preStop:
          exec:
            command: ["sleep", "30"]
    
  • Reduce controller sync interval: Modify your ALB Controller Deployment to set --sync-period=60s (down from the default 300s). This makes the controller check for Target Group updates more frequently, cutting down on registration/deregistration delays.

4. Long-Term Fix: Upgrade to ALB Ingress Controller 2.x

This is the most reliable solution to resolve all your issues:

  • 2.x versions are fully compatible with EKS 1.20+ and handle Pod readiness conditions, Target Group sync, and cross-namespace resources seamlessly.
  • It includes native optimizations for zero-downtime deployments, with simpler ReadinessGate configuration (or even auto-handling in some cases).
  • When upgrading:
    1. Uninstall the 1.1.4 controller completely.
    2. Install the latest 2.x version, following standard setup procedures for your cluster.
    3. Update your Ingress resources to use networking.k8s.io/v1 instead of extensions/v1beta1, and adjust any annotations that changed between versions.

内容的提问来源于stack exchange,提问作者Afik.A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 11:02:28