Istio Ingress配置AWS NLB IP模式TCP连接超时问题咨询
问题背景
近期在集群部署对接NLB的Istio Ingress,通过Service Annotation完成NLB与Istio Ingress的关联配置,所有路由规则均已配置完成。在集群后端部署测试服务器后,通过80端口HTTP协议向NLB发起curl请求开展连通性测试:
- 负载均衡类型设置为
nlb时,使用如下Annotation配置,curl测试可正常通过:
service.beta.kubernetes.io/aws-load-balancer-type: "nlb" service.beta.kubernetes.io/aws-load-balancer-security-groups: "<some security group>" service.beta.kubernetes.io/aws-load-balancer-cross-zone-load-balancing-enabled: "true" service.beta.kubernetes.io/aws-load-balancer-internal: "false" service.beta.kubernetes.io/aws-load-balancer-manage-backend-security-group-rules: "false"
- 将负载均衡类型改为
nlb-ip开启IP模式时,使用如下Annotation配置,curl请求在TCP连接建立阶段直接超时:
service.beta.kubernetes.io/aws-load-balancer-type: "nlb-ip" service.beta.kubernetes.io/aws-load-balancer-security-groups: "<some security group>" service.beta.kubernetes.io/aws-load-balancer-cross-zone-load-balancing-enabled: "true" service.beta.kubernetes.io/aws-load-balancer-internal: "false" service.beta.kubernetes.io/aws-load-balancer-manage-backend-security-group-rules: "false"
核心疑问:开启NLB的nlb-ip模式时是否需要配置特殊参数?
原因说明与配置要求
nlb实例模式和nlb-ipIP模式的流量转发逻辑完全不同,TCP连接超时是IP模式下缺失必要配置导致的,必须补充以下配置:
- 开启目标组客户端IP保留
实例模式下流量先转发到Node节点,经kube-proxy做SNAT后再转发到后端Pod;IP模式下NLB直接将流量转发到Pod IP,不经过Node上的kube-proxy链路,必须添加Annotationservice.beta.kubernetes.io/aws-load-balancer-target-group-attributes: "preserve_client_ip.enabled=true"。如果不开启该配置,Istio Ingress的Sidecar无法识别流量来源会直接丢弃连接,触发TCP超时。 - 调整安全组规则放通直连流量
当前配置了service.beta.kubernetes.io/aws-load-balancer-manage-backend-security-group-rules: "false",关闭了安全组规则自动管理。实例模式下仅需放通NLB到Node节点的安全组规则即可,IP模式下NLB直接和Pod IP通信,需要手动在Istio Ingress Pod所属安全组添加入方向规则,放通NLB所在网段到业务端口(如80、443)的TCP流量。 - 确认Service的externalTrafficPolicy配置
IP模式下不要将Istio Ingress对应的Service的externalTrafficPolicy设置为Local,需保持为Cluster,否则会因为流量未经过Node上的kube-proxy做链路转换,出现路由不可达问题。 - 移除不必要的代理协议配置
如果之前为实例模式开启了NLB Proxy Protocol v2配置,IP模式下需要关闭,否则Istio Ingress收到的报文格式不符合预期会直接断开连接,IP模式不需要额外开启该配置。
配置完成后重建NLB对应的Service,等待NLB目标组健康检查全部通过后,再发起curl测试即可恢复连通。
额外排查点:
nlb-ip模式下NLB目标组注册的是Pod IP,需要确认集群CNI插件支持Pod IP被VPC内的NLB直接路由访问。如果使用覆盖网络类CNI(比如默认配置的Flannel VXLAN),需要额外在VPC路由表添加Pod网段路由规则,否则NLB本身无法路由到Pod IP,同样会出现连接超时。
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

