Envoy新手求助:上游连接失败引发503错误,部署后20分钟正常后续异常
Hey there, sorry to hear you've been stuck with this Envoy issue for a week—let's break this down step by step to get to the bottom of it.
First, let's decode the error signals you're seeing to understand the core problem:
X-Envoy-Response-Code-Details: upstream_reset_before_response_started{connection_failure}: Envoy is failing to establish a TCP connection to your upstream service, or the connection is being reset before any response headers can be sent.X-Envoy-Response-Flags: UF,URX:UFstands for Upstream Connection Failure,URXfor Upstream Reset—this confirms the issue lives at the TCP connection layer between Envoy and your upstream, not an HTTP-level error from the service itself.
Your observation that problems start ~20 minutes after deployment is critical—it points to a gradual resource exhaustion or connection leak scenario, not an immediate misconfiguration. Let's walk through potential causes and actionable fixes:
1. Check Upstream Connection Resource Limits
Since direct access to your upstream works fine, the issue likely stems from Envoy maintaining a pool of persistent connections that eventually overwhelms the upstream's connection capacity.
- Monitor upstream TCP connections: Use tools like
ss -tulpnornetstaton your upstream servers to track active (ESTABLISHED) and idle (TIME_WAIT) connections. Look for a steady increase until they hit the process's file descriptor limit or system-level TCP max connections. - Check Envoy's connection stats: Query the Envoy admin endpoint (
curl localhost:9901/stats) and look for metrics like:cluster.consistent_cluster.upstream_cx_connect_failcluster.default_cluster.upstream_cx_connect_fail
A spike in these confirms connection failures are happening at scale.
2. Tune Envoy Connection Pool Configuration
Your current cluster config doesn't specify connection pool limits or idle timeouts, which can lead to unbound connection growth. Add these settings to both clusters:
clusters: - name: consistent_cluster # ... existing config ... circuit_breakers: thresholds: - priority: DEFAULT max_connections: 1000 # Adjust based on your upstream's maximum concurrent connections max_pending_requests: 500 connection_pool: tcp: idle_timeout: 30s # Clean up idle connections before the upstream closes them
circuit_breakersprevent Envoy from overwhelming the upstream with too many concurrent connections.idle_timeoutensures Envoy discards stale connections that the upstream may have already closed, avoiding reset errors when reusing them.
3. Adjust Health Check Sensitivity
Your health check has unhealthy_threshold: 1, meaning a single failed check marks the upstream as unhealthy. This is overly sensitive to network blips, especially with a 1-second interval. Tweak it to reduce false positives:
health_checks: - timeout: 1s interval: 1s unhealthy_threshold: 3 # Require 3 consecutive failures to mark as unhealthy healthy_threshold: 1 http_health_check: path: "/health"
- Also, check health check failure stats via the admin endpoint (
cluster.*.health_check.failure) to see if these align with your 503 errors. If health checks are failing unnecessarily, it could be causing Envoy to drop valid upstream instances.
4. Verify DNS Resolution Behavior
You're using STRICT_DNS for cluster discovery, which caches DNS results. If your upstream service's IP changes (e.g., Kubernetes pod restarts), Envoy may still try to connect to old IPs until the DNS cache refreshes.
- Check DNS refresh stats:
cluster.*.dns.refresh_failandcluster.*.dns.updatesto see if DNS resolution is working as expected. - Consider switching to
LOGICAL_DNSif your upstream IPs change frequently—it resolves DNS for each new connection instead of relying on cached results (note: this has minor performance tradeoffs).
5. Tweak Retry Policy to Avoid Amplifying Issues
Your retry policy includes connect-failure, which means Envoy will retry connections that fail. While this sounds helpful, it can flood the upstream with retries when connections start failing, making the problem worse. Add a backoff to slow down retries:
retry_policy: # ... existing config ... retry_back_off: base_interval: 0.1s max_interval: 1s
This ensures retries are spaced out, reducing load on the upstream during connection failures.
6. Enable Detailed Debug Logs
To get to the exact root cause, enable Envoy's debug logging (start Envoy with -l debug) and check the logs for connection-specific errors. You'll see messages like:
[debug] [connection] [source/common/network/connection_impl.cc:949] [C1234] closing socket: 104 (Connection reset by peer)
These logs will tell you exactly why the connection is failing (e.g., reset by upstream, timeout, refused).
内容的提问来源于stack exchange,提问作者sirsova

