GKE中负载均衡器无法连接服务的问题排查求助
问题:GKE集群Ingress返回502错误,NEG端点附加失败
我在GKE集群部署应用并配置Ingress负载均衡器,部署完成后Ingress已分配公网IP,但向该IP发送请求时返回502 Bad Gateway。查看Service事件发现NEG(Network Endpoint Group)端点附加失败。
应用配置
Deployment与Service配置
apiVersion: apps/v1 kind: Deployment metadata: name: api namespace: default spec: replicas: 1 selector: matchLabels: name: api template: metadata: labels: name: api spec: serviceAccountName: docker-sa containers: - name: api image: zhaoyi0113/rancher-go-api ports: - containerPort: 8080 apiVersion: v1 kind: Service metadata: name: api annotations: cloud.google.com/neg: '{"ingress": true}' spec: selector: name: api ports: - port: 80 targetPort: 8080 protocol: TCP type: NodePort
Ingress配置
apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: sidecar namespace: default spec: defaultBackend: service: name: api port: number: 80
错误现象
- Ingress状态正常但请求返回502:
$ kubectl get ingress NAME CLASS HOSTS ADDRESS PORTS AGE sidecar <none> * 107.178.245.193 80 28m
$ curl -i http://107.178.245.193/health HTTP/1.1 502 Bad Gateway Content-Type: text/html; charset=UTF-8 Referrer-Policy: no-referrer Content-Length: 332 Date: Tue, 16 Aug 2022 10:40:31 GMT <html><head> <meta http-equiv="content-type" content="text/html;charset=utf-8"> <title>502 Server Error</title> </head> <body text=#000000 bgcolor=#ffffff> <h1>Error: Server Error</h1> <h2>The server encountered a temporary error and could not complete your request.<p>Please try again in 30 seconds.</h2> <h2></h2> </body></html>
- Service事件显示NEG附加失败:
$ kubectl describe service api ... Events: Type Reason Age From Message ---- ------ ---- ---- ------- Warning AttachFailed 7s neg-controller Failed to Attach 2 network endpoint(s) (NEG "k8s1-29362abf-default-api-80-f2f1248a" in zone "australia-southeast2-a"): googleapi: Error 400: Invalid value for field 'resource.ipAddress': '10.0.1.18'. Specified IP address 10.0.1.18 doesn't belong to the (sub)network default or to the instance gke-gcp-cqrs-gcp-cqrs-node-pool-6b30ca5c-41q8., invalid Warning RetryFailed 7s neg-controller Failed to retry NEG sync for "default/api-k8s1-29362abf-default-api-80-f2f1248a--/80-8080-GCE_VM_IP_PORT-L7": maximum retry exceeded
问题根源与解决方案
根源分析
错误信息明确指出:Specified IP address 10.0.1.18 doesn't belong to the (sub)network default or to the instance,说明Pod的IP不在集群对应的VPC子网范围内,或者NEG控制器无法将该Pod IP关联到对应的GKE节点。常见触发场景:
- 集群子网配置变更,导致新分配的Pod IP不在原NEG关联的子网中
- Pod所在节点的网络配置异常,节点无法正确上报Pod所属子网信息
- NEG控制器缓存的子网信息与实际集群子网不匹配
解决方案
1. 验证Pod与节点的网络归属
先确认Pod的实际IP和所在节点的子网是否匹配:
# 获取Pod所在节点及Pod IP kubectl get pod -l name=api -o wide # 查看节点的子网IP段 gcloud compute instances describe <节点名称> --zone <节点所在区> | grep -A 5 networkIP
确认Pod IP是否属于节点所在子网的IP范围。
2. 重建NEG关联
删除并重新添加Service的NEG注解,触发GKE重新创建NEG并关联正确端点:
# 移除NEG注解 kubectl annotate service api cloud.google.com/neg- # 等待2-3分钟后重新添加注解 kubectl annotate service api cloud.google.com/neg='{"ingress": true}'
3. 重启Pod或节点(若上述方法无效)
- 重启异常Pod,让其重新分配IP:
kubectl delete pod -l name=api
- 如果节点网络配置异常,重启对应GKE节点:
gcloud compute instances restart <节点名称> --zone <节点所在区>
4. 检查集群子网配置
确认集群使用的VPC子网未被修改或删除:
gcloud container clusters describe <集群名称> --zone <集群所在区> | grep -A 10 "network"
确保子网的IP范围包含Pod的IP段。
内容的提问来源于stack exchange,提问作者Joey Yi Zhao
相关产品推荐
相关产品推荐

