Istio异常检测导致负载均衡失效且无触发指标的问题排查求助
First off, nice work narrowing down that Outlier Detection (OD) is the root cause—removing it restoring load balance is a key clue. Let’s dig into why this might be happening even without seeing trigger metrics, and how to get to the bottom of it.
1. Your Configuration Red Flags
Looking at your updated OD config, even with both error thresholds set to 0, some subtle behavior in Istio 1.10 could be causing unintended ejections:
outlierDetection: interval: 10s maxEjectionPercent: 50
- Empty threshold values: While setting
consecutive5xxErrorsandconsecutiveGatewayErrorsto 0 disables those error checks, Istio’s OD might still have implicit checks enabled (like slow requests, though you haven’t configured them) or there could be a bug in 1.10 where incomplete OD configs lead to unexpected instance marking. - Connection Pool limits: Your
connectionPoolsettings (maxConnections: 500,http1MaxPendingRequests: 1000) might be interacting with OD. If a pod hits these limits, Envoy might mark it as unhealthy without triggering standard 503/UO metrics—especially if backpressure leads to connection timeouts instead of explicit errors.
2. Why You’re Not Seeing Trigger Metrics
Istio 1.10 has a few quirks with OD metrics that could explain the missing data:
- Wrong metric focus: The
envoy_cluster_circuit_breakers_default_cx_openmetric tracks circuit breakers, not OD ejections. For OD, you should targetenvoy_cluster_outlier_detection_ejections_active(current live ejections) orenvoy_cluster_outlier_detection_ejections_total(total ejections over time). - Prometheus scraping gaps: Sidecars expose metrics on port 15090 by default, not 15000. Double-check your Prometheus scrape job targets this port and includes pods with the
istio-prometheus-merge: "true"annotation. - UO flags only trigger on explicit rejects: The
response_flags="UO"(Unhealthy Origin) only appears when Envoy actively rejects a request to an ejected pod. If the load balancer is just avoiding the pod entirely without rejecting requests, this flag won’t show up.
3. Step-by-Step Troubleshooting
Let’s get concrete data to confirm what’s happening:
A. Inspect Envoy’s Applied Cluster Config
Run this on a client pod (the one sending traffic to some-service) to see the actual OD logic Istio pushed to Envoy:
istioctl pc cluster <client-pod-name> -n <client-namespace> -o yaml | grep -A 20 "outlier_detection"
Look for:
ejection_thresholdvalues matching your config- Unexpected
slow_request_thresholdorconsecutive_slow_requestssettings (Istio might apply implicit defaults even if you didn’t configure them) max_ejection_percentcorrectly set to 50
B. Check Real-Time Ejections via Envoy Admin
Access the Envoy admin interface directly to see if pods are being ejected right now:
kubectl exec -it <client-pod-name> -n <client-namespace> -- curl http://localhost:15000/clusters | grep -E "(some-service|outlier_detection)"
Look for lines like:
some-service.some-namespace.svc.cluster.local|80::outlier_detection_ejections_active: 1
This will confirm if a pod is marked unhealthy even if metrics aren’t capturing it.
C. Validate Prometheus Metric Presence
Run these queries in Prometheus to check if OD metrics exist:
sum(envoy_cluster_outlier_detection_ejections_total) by (destination_cluster) sum(envoy_cluster_outlier_detection_ejections_active) by (destination_cluster)
If no results return, fix your Prometheus scrape job to target sidecar port 15090 with path /metrics.
D. Test with Fully Explicit OD Config
Instead of leaving threshold fields out, try explicitly disabling all checks to avoid implicit defaults:
outlierDetection: consecutive5xxErrors: 0 consecutiveGatewayErrors: 0 consecutiveLocalOriginFailures: 0 interval: 10s baseEjectionTime: 30s maxEjectionPercent: 50
If load balance works after this, the incomplete config was likely the issue.
4. Additional Istio 1.10-Specific Checks
- Known bugs: Istio 1.10 had several OD-related bugs (like incorrect ejection logic for gateway errors). Consider upgrading to a newer 1.10 patch version or to 1.11+ if possible, as many of these issues were fixed.
- Retry policy interactions: Your VirtualService’s 5 retry attempts could amplify connection pool pressure. If retries hit the same pod repeatedly, it might trigger OD even without explicit errors. Try reducing retry attempts temporarily to test.
内容的提问来源于stack exchange,提问作者Asuu

