Grafana Prometheus默认查询错误及多Envoy指标求和方法咨询
Hey there! Let's break down your two questions and walk through practical solutions for each.
envoy_cluster_cluster_serviceN_upstream_rq_time Metrics Prometheus has flexible pattern-matching tools that make this straightforward. Here are two reliable approaches:
Approach 1: Match by metric name pattern
Use a wildcard .* to target all service number suffixes, then wrap the match in the sum() aggregation function:
sum(envoy_cluster_cluster_service.*_upstream_rq_time)
This will directly sum every metric that follows the envoy_cluster_cluster_service[0-9]+_upstream_rq_time naming pattern (covering service1 through service100).
Approach 2: Match by label (more robust long-term)
If your Envoy metrics include a cluster label with values like cluster_service1 or cluster_service2, using label regex matching is more maintainable (it avoids breaking if metric naming conventions change later):
sum(envoy_cluster_upstream_rq_time{cluster=~"cluster_service[0-9]+"})
- The
=~operator enables regex pattern matching [0-9]+ensures we only target services with numeric suffixes, preventing accidental matches with unrelated clusters
You can also refine the result with without or by clauses to preserve specific labels. For example, to sum metrics while keeping the envoy_version label:
sum without (cluster) (envoy_cluster_upstream_rq_time{cluster=~"cluster_service[0-9]+"})
Since you didn't share the exact error message, here are the most common fixes for default query issues:
- Check Prometheus data source connectivity: Head to Grafana's Data Sources > Prometheus, click "Test Connection" to confirm Grafana can reach your Prometheus instance. Double-check the URL, port, and any firewall rules blocking traffic between the two tools.
- Validate the query directly in Prometheus: Copy the problematic default query from Grafana, paste it into Prometheus's Expression Browser (typically at
http://<prometheus-ip>:9090/graph). If it returns no results, the metric might not be scraped correctly, or the query has a typo. - Adjust the time range: Grafana's default time range (like "Last 6 hours") might not cover periods where your metrics have data. Try expanding it to "Last 24 hours" or a custom range that aligns with when your services were active.
- Verify authentication settings: If your Prometheus instance uses basic auth or API tokens, make sure you've entered the correct credentials in Grafana's data source configuration.
- Confirm Prometheus scrape configs: Ensure your Prometheus
scrape_configsinclude the Envoy metrics endpoint, and that the endpoint returns valid metrics (test this by visiting the endpoint in a browser or usingcurl).
内容的提问来源于stack exchange,提问作者kosnkov

