GKE部署的Go前端网站间歇性返回502 Bad Gateway求助排查
问题描述
我用Go开发了前端站点,通过Docker部署到GKE。访问http://staging.crcl-app.com、http://www.staging.crcl-app.com时,网站有时能正常加载,有时返回502 Bad Gateway。我判断问题大概率不在部署环节,可能和Ingress有关?已配置Cloud DNS将域名指向ingress.yaml中定义的“crcl-global”静态IP,相关配置截图已附上。
相关配置文件
deployment.yaml
apiVersion: apps/v1 kind: Deployment metadata: name: staging-deployment namespace: admin-ns labels: app: staging spec: replicas: 1 selector: matchLabels: app: staging tier: web template: metadata: labels: app: staging tier: web spec: containers: - name: staging image: # removed for stackoverflow ports: - containerPort: 3001 readinessProbe: httpGet: path: /healthz # GKE readinessProbe healthcheck port: 3001 initialDelaySeconds: 5 periodSeconds: 10 livenessProbe: httpGet: path: /healthz # GKE livenessProbe healthcheck port: 3001 initialDelaySeconds: 15 periodSeconds: 20 resources: limits: memory: 512Mi cpu: "1" requests: memory: 256Mi cpu: "0.2" serviceAccountName: "admin-sa"
ingress.yaml
apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: crcl-staging-ingress # name of service namespace: admin-ns # this is the role in GCP it has, we need to give it perms so it can access secret manager, storage, etc. annotations: spec.ingressClassName: "nginx" nginx.ingress.kubernetes.io/use-regex: "true" nginx.ingress.kubernetes.io/rewrite-target: /$1 kubernetes.io/ingress.global-static-ip-name: crcl-global # Expose IP networking.gke.io/managed-certificates: crcl-app-cert # Google managed certificates labels: app: staging spec: defaultBackend: service: name: crcl-gke-backend port: number: 80 rules: - host: www.staging.crcl-app.com http: paths: - path: / pathType: ImplementationSpecific backend: service: name: crcl-gke-backend port: number: 80 - host: api.staging.crcl-app.com http: paths: - path: /api/* pathType: ImplementationSpecific backend: service: name: crcl-gke-backend port: number: 80 - host: staging.crcl-app.com http: paths: - path: / pathType: ImplementationSpecific backend: service: name: crcl-gke-backend port: number: 80 --- apiVersion: v1 kind: Service metadata: name: crcl-gke-backend namespace: admin-ns labels: app: staging spec: type: NodePort # Internal staging IP (Routed to by ingress) selector: app: staging tier: web ports: - port: 80 targetPort: 3001
排查建议
针对间歇性502问题,从Ingress及关联组件入手,推荐做以下排查:
- 检查Ingress Controller稳定性:确认Nginx Ingress Controller的Pod无重启、无异常日志,执行命令:
kubectl get pods -n ingress-nginx kubectl logs -n ingress-nginx <ingress-controller-pod-name> - 验证Service后端端点状态:查看
crcl-gke-backend是否正确关联Pod,端点是否正常,执行:
若端点异常,大概率是Pod就绪探针失败,导致Service停止向该Pod转发流量。kubectl get endpoints crcl-gke-backend -n admin-ns - 检查Pod探针状态:查看
staging-deployment的Pod是否有重启记录,就绪/存活探针是否失败,执行:
探针失败会让Pod暂时被踢出Service端点列表,Ingress转发流量时就会返回502。kubectl get pods -n admin-ns kubectl describe pod <staging-pod-name> -n admin-ns - 修正Ingress重写规则:当前全局配置了
use-regex: "true"和rewrite-target: /$1,但主域名路径为/,正则匹配下会导致重写异常。建议移除全局重写配置,或仅针对api.staging.crcl-app.com单独配置重写规则,避免影响主站流量。 - 检查GKE托管证书状态:确认
crcl-app-cert是否成功颁发,证书状态异常会导致Ingress流量中断,执行:kubectl get managedcertificates crcl-app-cert -n admin-ns - 确认静态IP绑定状态:在GCP控制台查看
crcl-global静态IP是否正确绑定到Ingress,未被其他资源占用。
内容的提问来源于stack exchange,提问作者Matthew Masody
相关产品推荐
相关产品推荐

