You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置Envoy Sidecar实现gRPC上游服务故障转移?

Fixing Envoy Sidecar Failures When a gRPC Server Goes Down

Hey there! I see you're new to Envoy Proxy and running into an issue where your gRPC client throws errors when one backend server goes down—even with ROUND_ROBIN load balancing enabled. Let's break down why this is happening and fix your configuration step by step.

Why This Is Happening

Right now, your Envoy cluster uses a static type with ROUND_ROBIN, but you haven’t set up health checks. Envoy has no way of knowing when a gRPC server goes offline, so it keeps routing requests to the dead instance. That’s exactly why you’re seeing those connection failure errors in both your client logs and Envoy’s debug output.

The Solution: Add gRPC Health Checks & Retry Policies

To make Envoy detect unhealthy backends and stop sending traffic to them, we need to add gRPC-specific health checks to your cluster. We’ll also add a retry policy to handle transient failures gracefully.

Here’s your updated Envoy configuration with all necessary fixes:

admin:
  access_log_path: "/tmp/admin_access.log"
  address:
    socket_address:
      address: "10.19.17.188"
      port_value: 12000
static_resources:
  listeners:
  - name: "grpc-listener"
    address:
      socket_address:
        address: "10.19.17.188"
        port_value: 12001
    filter_chains:
    - filters:
      - name: "envoy.http_connection_manager"
        config:
          stat_prefix: "ingress"
          codec_type: "AUTO"
          route_config:
            name: "grpc-route"
            virtual_hosts:
            - name: "grpc-route"
              domains:
              - "*"
              routes:
              - match:
                  prefix: "/"
                route:
                  cluster: "grpc-service"
                  # Add retry policy for gRPC connection failures
                  retry_policy:
                    retry_on: "connect-failure,refused-stream,unavailable"
                    num_retries: 2
                    per_try_timeout: "0.5s"
          http_filters:
          - name: "envoy.router"
  clusters:
  - name: "grpc-service"
    connect_timeout: "0.25s"
    type: "static"
    lb_policy: "ROUND_ROBIN"
    http2_protocol_options: {}
    # Add gRPC health check configuration
    health_checks:
    - timeout: "0.5s"
      interval: "5s"
      unhealthy_threshold: 2
      healthy_threshold: 2
      grpc_health_check:
        service_name: "" # Leave empty if using default gRPC health service
    hosts:
    - socket_address:
        address: "10.19.17.188"
        port_value: 12011
    - socket_address:
        address: "10.19.17.188"
        port_value: 12012

Key Configuration Breakdown

Let’s walk through the critical additions:

  1. gRPC Health Checks

    • grpc_health_check: Uses the standard gRPC Health Checking Protocol (supported by most gRPC frameworks out of the box) to probe backend health.
    • interval: "5s": How often Envoy sends health check requests to each backend.
    • timeout: "0.5s": How long Envoy waits for a response before marking an instance as unhealthy.
    • unhealthy_threshold: 2: Number of consecutive failed checks needed to remove an instance from the load pool.
    • healthy_threshold: 2: Number of consecutive successful checks needed to add a recovered instance back to the pool.
  2. Retry Policy

    • retry_on: "connect-failure,refused-stream,unavailable": Triggers retries for common gRPC errors related to dead or unreachable backends.
    • num_retries: 2: Maximum number of retries per failed request.
    • per_try_timeout: "0.5s": Timeout for each individual retry attempt to avoid hanging.

Quick Notes

  • Ensure your gRPC servers have the health checking service enabled. Most frameworks (Python, Go, Java) include this by default—you just need to register the health service with your server.
  • If your server uses a custom health service name, replace the empty string in service_name with your actual service identifier.

With these changes, Envoy will automatically detect offline backends, stop routing traffic to them, and retry failed requests on healthy instances. Your gRPC client should no longer throw errors when one server goes down.

内容的提问来源于stack exchange,提问作者Ian.Zhang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:05:36