Azure私有TKG集群用Helm部署istio-ingressgateway失败:无法获取现有内部LB IP
Istio IngressGateway LoadBalancer服务在Azure TKG集群中Pending的排查与解决
问题背景
在Azure私有网络中部署的TKG 2.1.1(Kubernetes 1.24.10)集群(已由Tanzu预配内部负载均衡器),通过Helm部署Istio IngressGateway时,服务的EXTERNAL-IP始终处于<pending>状态,Helm状态显示failed但提示安装成功。
部署配置与现象
安装命令:
helm install -f values.yaml istio-ingressgateway istio/gateway -n istio-ingress --wait
尝试过的values.yaml配置包括:
- 基础内部LB配置
service: type: LoadBalancer ports: - name: status-port port: 15021 protocol: TCP targetPort: 15021 - name: http2 port: 80 protocol: TCP targetPort: 80 - name: https port: 443 protocol: TCP targetPort: 443 annotations: service.beta.kubernetes.io/azure-load-balancer-internal: 'true'
- 指定现有LB IP
service: type: LoadBalancer ports: - name: status-port port: 15021 protocol: TCP targetPort: 15021 - name: http2 port: 80 protocol: TCP targetPort: 80 - name: https port: 443 protocol: TCP targetPort: 443 annotations: service.beta.kubernetes.io/azure-load-balancer-internal: 'true' service.beta.kubernetes.io/azure-load-balancer-ipv4: <existing lb ip>
- 指定内部子网
service: type: LoadBalancer ports: - name: status-port port: 15021 protocol: TCP targetPort: 15021 - name: http2 port: 80 protocol: TCP targetPort: 80 - name: https port: 443 protocol: TCP targetPort: 443 annotations: service.beta.kubernetes.io/azure-load-balancer-internal: 'true' service.beta.kubernetes.io/azure-load-balancer-internal-subnet: app-pln-snet
服务核心状态:
kubectl get service istio-ingressgateway -n istio-ingress -o wide NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE SELECTOR istio-ingressgateway LoadBalancer 100.69.48.176 <pending> 15021:32090/TCP,80:31815/TCP,443:30364/TCP 42m app=istio-ingressgateway,istio=ingressgateway
排查与解决步骤
1. 验证Azure云控制器管理器(CCM)状态
Azure上的Kubernetes集群依赖CCM处理LoadBalancer服务的创建与同步,先确认CCM组件正常:
- 检查CCM Pod运行状态:
kubectl get pods -n kube-system | grep cloud-controller-manager - 如果Pod异常,查看日志定位错误:
kubectl logs -n kube-system <cloud-controller-manager-pod-name> - 确认TKG部署时使用的服务账号拥有Azure资源权限,至少需要
Network Contributor角色,用于管理负载均衡器和网络资源。
2. 检查子网与IP配置的有效性
- 确认
service.beta.kubernetes.io/azure-load-balancer-internal-subnet指定的app-pln-snet子网存在于集群所在的VNet中,且子网区域与集群节点一致。 - 检查子网的IP地址池是否有剩余可用IP,若子网IP耗尽,LB无法分配地址。
- 如果指定了
service.beta.kubernetes.io/azure-load-balancer-ipv4,确保该IP是app-pln-snet子网内的未分配静态IP,且未被其他Azure资源占用。
3. 确认Istio Gateway Pod状态
虽然Endpoints已存在,但仍需验证Gateway Pod是否正常运行:
- 查看Pod状态:
kubectl get pods -n istio-ingress - 检查Pod日志是否有网络初始化或连接错误:
kubectl logs -n istio-ingress <istio-ingressgateway-pod-name>
4. 触发LoadBalancer服务重新同步
如果CCM状态正常,可通过以下方式触发服务重新处理:
- 卸载并重新部署IngressGateway:
helm uninstall istio-ingressgateway -n istio-ingress kubectl delete namespace istio-ingress --ignore-not-found kubectl create namespace istio-ingress helm install -f values.yaml istio-ingressgateway istio/gateway -n istio-ingress - 或通过修改注解触发CCM重新同步:
kubectl annotate service istio-ingressgateway -n istio-ingress service.beta.kubernetes.io/azure-load-balancer-internal- kubectl annotate service istio-ingressgateway -n istio-ingress service.beta.kubernetes.io/azure-load-balancer-internal='true'
5. 复用现有Tanzu预配的内部LB
如果需要连接到已有的内部LB,需添加额外注解指定LB的资源信息:
service: type: LoadBalancer ports: - name: status-port port: 15021 protocol: TCP targetPort: 15021 - name: http2 port: 80 protocol: TCP targetPort: 80 - name: https port: 443 protocol: TCP targetPort: 443 annotations: service.beta.kubernetes.io/azure-load-balancer-internal: 'true' service.beta.kubernetes.io/azure-load-balancer-resource-group: <lb-resource-group> service.beta.kubernetes.io/azure-load-balancer-name: <existing-lb-name>
确保现有LB的资源组、名称正确,且LB与集群处于同一VNet中。
内容的提问来源于stack exchange,提问作者Chad Larkin
相关产品推荐
相关产品推荐

