如何配置搭载Envoy与gRPC-Web的GKE Autopilot及解决健康检查失败问题
核心问题定位
- GCP负载均衡的健康检查默认不会携带你Envoy配置中要求的
x-envoy-livenessprobe请求头:你在Envoy的健康检查过滤器中配置了必须同时匹配/healthz路径和x-envoy-livenessprobe: healthz头才会返回健康状态,但BackendConfig中没有配置该自定义头,导致健康检查请求被Envoy拒绝,返回非200状态码。 - 你看到的就绪探针显示
http-get HTTP://:0/healthz属于Kubernetes的显示异常,实际运行时会自动映射到你定义的名为http的8080端口,不是配置错误。 - Envoy配置中上游gRPC服务的地址写为
0.0.0.0虽然不影响同Pod内通信,但规范写法应为127.0.0.1。
修复步骤
- 修改BackendConfig配置,为健康检查添加自定义请求头,修改后的配置如下:
apiVersion: cloud.google.com/v1 kind: BackendConfig metadata: name: grammar-games-bec annotations: cloud.google.com/neg: '{"ingress": true}' spec: sessionAffinity: affinityType: "CLIENT_IP" healthCheck: checkIntervalSec: 15 port: 8080 type: HTTP requestPath: /healthz # 新增自定义请求头配置 customHeaders: request: x-envoy-livenessprobe: "healthz" timeoutSec: 60
- 可选优化:修改Envoy ConfigMap中上游集群的地址为
127.0.0.1,避免意外绑定:
clusters: - name: grammar-games-core-grpc connect_timeout: 0.5s type: logical_dns lb_policy: ROUND_ROBIN http2_protocol_options: {} load_assignment: cluster_name: grammar-games-core-grpc endpoints: - lb_endpoints: - endpoint: address: socket_address: address: 127.0.0.1 port_value: 52001 health_checks: timeout: 1s interval: 10s unhealthy_threshold: 2 healthy_threshold: 2 grpc_health_check: {}
- 重新应用修改后的配置,等待GKE Ingress资源完全生效(通常需要5-10分钟),健康检查即可恢复正常。
- 可通过
kubectl exec进入Pod手动验证健康检查接口是否正常:curl -H "x-envoy-livenessprobe: healthz" http://127.0.0.1:8080/healthz
正常情况下会返回200状态码。
内容的提问来源于stack exchange,提问作者Renee Revis
相关产品推荐
相关产品推荐

