You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GCE Ingress关联GRPC服务后端不健康状态排查求助

问题:GCE Ingress 后端服务持续不健康,Pod探针却正常

我正尝试通过gce-ingress访问gRPC服务端,部署完成后Ingress始终返回"Some backend services are in UNHEALTHY state"。所有Pod运行无报错,手动通过kubectl exec进入Pod执行/bin/grpc_health_probe -addr=:8000能得到"SERVING"结果,但Ingress关联的后端服务一直显示不健康状态。


相关配置文件

Ingress 配置(Ingress.yaml)

apiVersion: cloud.google.com/v1
kind: BackendConfig
metadata:
  name: app-backend-config
spec:
  customRequestHeaders:
    headers:
    - "TE:trailers"
---

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: ingress-prod
  annotations:
    kubernetes.io/ingress.global-static-ip-name: dummy-ingress
    kubernetes.io/ingress.allow-http: "false"
    cert-manager.io/issuer: issuer
    cloud.google.com/backend-config: '{"default": "app-backend-config"}'
  labels:
    name: ingress-app
spec:
  tls:
  - hosts:
    - domain.name
    secretName: secret-tls
  rules:
  - host: domain.name
    http:
      paths:
      - path: /*
        pathType: ImplementationSpecific
        backend:
          service:
            name: app-server-headless
            port:
              number: 8000

应用服务端部署配置

apiVersion: apps/v1
kind: Deployment
metadata:
  name: dummy-app-server

  labels:
    app: app-server
spec:
  replicas: 1
  selector:
    matchLabels:
      app: app-server
  template:
    metadata:
      labels:
        app: app-server
    spec:
      containers:
        name: app-server
        image: gcr.io/emeritus-data-science/image:latest
        command: ["python3" , "/var/app/api_server/main.py"]
        imagePullPolicy: Always
        resources: # limit the resources
          requests:
            memory: 1Gi
            cpu: "1"
          limits:
            memory: 2Gi
            cpu: "1"
        volumeMounts:
        - mountPath: /secrets/gcloud-auth
          name: gcloud-auth
          readOnly: true
        ports:
        - containerPort: 8000
        readinessProbe:
          exec:
            command: [ "/bin/grpc_health_probe", "-addr=:8000" ]
          initialDelaySeconds: 30
          timeoutSeconds: 5
          periodSeconds: 10
          failureThreshold: 2
        livenessProbe:
          exec:
            command: [ "/bin/grpc_health_probe", "-addr=:8000" ]
          initialDelaySeconds: 60
          timeoutSeconds: 5
          periodSeconds: 10
          failureThreshold: 2

      volumes:
      - name: gcloud-auth
        secret:
          secretName: gcloud
---
apiVersion: v1
kind: Service
metadata:
  name: app-server-headless
  annotations:
    cloud.google.com/app-protocols: '{"grpc":"HTTP2"}'
spec:
  type: ClusterIP
  clusterIP: None
  selector:
    app: app-server
  ports:
    - protocol: TCP
      port: 8000
      targetPort: 8000
      name: grpc

gRPC服务端代码片段

server = grpc.server(futures.ThreadPoolExecutor(max_workers=10))
master_pb2_grpc.add_EventBusOneofServiceServicer_to_server(
    EventBusServiceServicer(), server
)

health_pb2_grpc.add_HealthServicer_to_server(health.HealthServicer(), server)

server.add_insecure_port("0.0.0.0:8000")
server.start()
LOG.info("server started")

def handle_sigterm(*_):
    print("Received shutdown signal")
    all_rpcs_done_event = server.stop(30)
    all_rpcs_done_event.wait(30)
    print("Shut down gracefully")

signal(SIGTERM, handle_sigterm)
server.wait_for_termination()

当前后端服务健康检查配置

Description: Default kubernetes L7 Loadbalancing health check for NEG.
Path: /
Protocol: HTTP/2
Port specification: Serving port
Proxy protocol: NONE
Logs: Disabled
Interval: 15 seconds
Timeout: 15 seconds
Healthy threshold: 1 success
Unhealthy threshold: 2 consecutive failures


疑问

是就绪探针和存活探针配置存在问题,还是部署中的其他环节导致了该异常?

内容的提问来源于stack exchange,提问作者ak1234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 14:05:47