GCE Ingress关联GRPC服务后端不健康状态排查求助
问题:GCE Ingress 后端服务持续不健康,Pod探针却正常
我正尝试通过gce-ingress访问gRPC服务端,部署完成后Ingress始终返回"Some backend services are in UNHEALTHY state"。所有Pod运行无报错,手动通过kubectl exec进入Pod执行/bin/grpc_health_probe -addr=:8000能得到"SERVING"结果,但Ingress关联的后端服务一直显示不健康状态。
相关配置文件
Ingress 配置(Ingress.yaml)
apiVersion: cloud.google.com/v1 kind: BackendConfig metadata: name: app-backend-config spec: customRequestHeaders: headers: - "TE:trailers" --- apiVersion: networking.k8s.io/v1 kind: Ingress metadata: name: ingress-prod annotations: kubernetes.io/ingress.global-static-ip-name: dummy-ingress kubernetes.io/ingress.allow-http: "false" cert-manager.io/issuer: issuer cloud.google.com/backend-config: '{"default": "app-backend-config"}' labels: name: ingress-app spec: tls: - hosts: - domain.name secretName: secret-tls rules: - host: domain.name http: paths: - path: /* pathType: ImplementationSpecific backend: service: name: app-server-headless port: number: 8000
应用服务端部署配置
apiVersion: apps/v1 kind: Deployment metadata: name: dummy-app-server labels: app: app-server spec: replicas: 1 selector: matchLabels: app: app-server template: metadata: labels: app: app-server spec: containers: name: app-server image: gcr.io/emeritus-data-science/image:latest command: ["python3" , "/var/app/api_server/main.py"] imagePullPolicy: Always resources: # limit the resources requests: memory: 1Gi cpu: "1" limits: memory: 2Gi cpu: "1" volumeMounts: - mountPath: /secrets/gcloud-auth name: gcloud-auth readOnly: true ports: - containerPort: 8000 readinessProbe: exec: command: [ "/bin/grpc_health_probe", "-addr=:8000" ] initialDelaySeconds: 30 timeoutSeconds: 5 periodSeconds: 10 failureThreshold: 2 livenessProbe: exec: command: [ "/bin/grpc_health_probe", "-addr=:8000" ] initialDelaySeconds: 60 timeoutSeconds: 5 periodSeconds: 10 failureThreshold: 2 volumes: - name: gcloud-auth secret: secretName: gcloud --- apiVersion: v1 kind: Service metadata: name: app-server-headless annotations: cloud.google.com/app-protocols: '{"grpc":"HTTP2"}' spec: type: ClusterIP clusterIP: None selector: app: app-server ports: - protocol: TCP port: 8000 targetPort: 8000 name: grpc
gRPC服务端代码片段
server = grpc.server(futures.ThreadPoolExecutor(max_workers=10)) master_pb2_grpc.add_EventBusOneofServiceServicer_to_server( EventBusServiceServicer(), server ) health_pb2_grpc.add_HealthServicer_to_server(health.HealthServicer(), server) server.add_insecure_port("0.0.0.0:8000") server.start() LOG.info("server started") def handle_sigterm(*_): print("Received shutdown signal") all_rpcs_done_event = server.stop(30) all_rpcs_done_event.wait(30) print("Shut down gracefully") signal(SIGTERM, handle_sigterm) server.wait_for_termination()
当前后端服务健康检查配置
Description: Default kubernetes L7 Loadbalancing health check for NEG.
Path: /
Protocol: HTTP/2
Port specification: Serving port
Proxy protocol: NONE
Logs: Disabled
Interval: 15 seconds
Timeout: 15 seconds
Healthy threshold: 1 success
Unhealthy threshold: 2 consecutive failures
疑问
是就绪探针和存活探针配置存在问题,还是部署中的其他环节导致了该异常?
内容的提问来源于stack exchange,提问作者ak1234
相关产品推荐
相关产品推荐

