Prometheus Summary分位数的PromQL用法:如何获取端点整体p95延迟
Got it, let's work through this together! Since you're using a client-side precomputed Summary metric, you're already ahead on query speed—now let's make sure you're pulling the right overall p95 latency for your endpoint.
First, let's recap how your Summary metric is structured in Prometheus: when your client reports my_metric_api_latency_seconds with precomputed percentiles, you'll have time series tagged with quantile="0.95" (alongside the p50/p90/p99 ones), plus the _sum and _count metrics for total latency and request volume.
Here's how to query based on your needs:
1. Get the current p95 latency for your endpoint
If your endpoint is identified by a tag (like endpoint="/api/checkout"), just filter for the p95 quantile and your target endpoint:
my_metric_api_latency_seconds{quantile="0.95", endpoint="/your-target-endpoint"}
This pulls the precomputed value directly, which is why it's fast—no on-the-fly percentile calculation needed.
2. Aggregate across dimensions (with a critical caveat)
If your metric has extra tags (like instance="web-01" or region="eu-west"), and you want an overall p95 across those dimensions, keep this in mind:
Important: Percentiles are not statistically aggregatable. Averaging or summing p95 values from different subsets of traffic will not give you the true overall p95 of all requests to the endpoint.
That said, if your use case can tolerate an approximate value (e.g., all instances handle similar traffic volumes), you can use avg with without to drop the extra tags:
avg(my_metric_api_latency_seconds{quantile="0.95", endpoint="/your-target-endpoint"}) without (instance, region)
Replace instance, region with the tags you want to exclude from the aggregation.
For an accurate overall percentile (if you need it)
To get a true overall p95 across all requests, the client-side precomputed percentiles won't help—you need raw latency data to aggregate. If you can switch to a Histogram metric, that's the way to go because Histograms support accurate aggregated percentile calculations with histogram_quantile().
If you're stuck with the Summary metric, you can calculate average latency (though this isn't a percentile) using the _sum and _count fields:
sum(my_metric_api_latency_seconds_sum{endpoint="/your-target-endpoint"}) without (instance, region) / sum(my_metric_api_latency_seconds_count{endpoint="/your-target-endpoint"}) without (instance, region)
3. Track p95 latency over time
If you want to see how the p95 latency has trended over the last hour, use avg_over_time (or max_over_time/min_over_time depending on what you care about) to roll up the precomputed values:
avg_over_time(my_metric_api_latency_seconds{quantile="0.95", endpoint="/your-target-endpoint"}[1h])
This will show you the average p95 value over 1-hour windows, making it easy to spot trends.
内容的提问来源于stack exchange,提问作者sb2k18

