如何配置Envoy Sidecar实现gRPC上游服务故障转移?
Hey there! I see you're new to Envoy Proxy and running into an issue where your gRPC client throws errors when one backend server goes down—even with ROUND_ROBIN load balancing enabled. Let's break down why this is happening and fix your configuration step by step.
Why This Is Happening
Right now, your Envoy cluster uses a static type with ROUND_ROBIN, but you haven’t set up health checks. Envoy has no way of knowing when a gRPC server goes offline, so it keeps routing requests to the dead instance. That’s exactly why you’re seeing those connection failure errors in both your client logs and Envoy’s debug output.
The Solution: Add gRPC Health Checks & Retry Policies
To make Envoy detect unhealthy backends and stop sending traffic to them, we need to add gRPC-specific health checks to your cluster. We’ll also add a retry policy to handle transient failures gracefully.
Here’s your updated Envoy configuration with all necessary fixes:
admin: access_log_path: "/tmp/admin_access.log" address: socket_address: address: "10.19.17.188" port_value: 12000 static_resources: listeners: - name: "grpc-listener" address: socket_address: address: "10.19.17.188" port_value: 12001 filter_chains: - filters: - name: "envoy.http_connection_manager" config: stat_prefix: "ingress" codec_type: "AUTO" route_config: name: "grpc-route" virtual_hosts: - name: "grpc-route" domains: - "*" routes: - match: prefix: "/" route: cluster: "grpc-service" # Add retry policy for gRPC connection failures retry_policy: retry_on: "connect-failure,refused-stream,unavailable" num_retries: 2 per_try_timeout: "0.5s" http_filters: - name: "envoy.router" clusters: - name: "grpc-service" connect_timeout: "0.25s" type: "static" lb_policy: "ROUND_ROBIN" http2_protocol_options: {} # Add gRPC health check configuration health_checks: - timeout: "0.5s" interval: "5s" unhealthy_threshold: 2 healthy_threshold: 2 grpc_health_check: service_name: "" # Leave empty if using default gRPC health service hosts: - socket_address: address: "10.19.17.188" port_value: 12011 - socket_address: address: "10.19.17.188" port_value: 12012
Key Configuration Breakdown
Let’s walk through the critical additions:
gRPC Health Checks
grpc_health_check: Uses the standard gRPC Health Checking Protocol (supported by most gRPC frameworks out of the box) to probe backend health.interval: "5s": How often Envoy sends health check requests to each backend.timeout: "0.5s": How long Envoy waits for a response before marking an instance as unhealthy.unhealthy_threshold: 2: Number of consecutive failed checks needed to remove an instance from the load pool.healthy_threshold: 2: Number of consecutive successful checks needed to add a recovered instance back to the pool.
Retry Policy
retry_on: "connect-failure,refused-stream,unavailable": Triggers retries for common gRPC errors related to dead or unreachable backends.num_retries: 2: Maximum number of retries per failed request.per_try_timeout: "0.5s": Timeout for each individual retry attempt to avoid hanging.
Quick Notes
- Ensure your gRPC servers have the health checking service enabled. Most frameworks (Python, Go, Java) include this by default—you just need to register the health service with your server.
- If your server uses a custom health service name, replace the empty string in
service_namewith your actual service identifier.
With these changes, Envoy will automatically detect offline backends, stop routing traffic to them, and retry failed requests on healthy instances. Your gRPC client should no longer throw errors when one server goes down.
内容的提问来源于stack exchange,提问作者Ian.Zhang

