You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Prometheus慢查询调试:排查Grafana超时查询性能问题

Great question—let’s break this down step by step. Your hunch about high cardinality from multi-label metrics is likely on the mark, so let’s start with how to validate that, then dive into practical query performance analysis tricks.

一、验证多标签指标的样本量/基数问题

These steps will help you confirm if excessive time series from multi-label combinations are causing your timeouts:

  • Count total time series for the metric
    Run this PromQL query directly in Grafana or Prometheus:
    count({__name__="your_target_metric"})
    
    If the result is in the tens of thousands or higher, you’re dealing with high cardinality. Even queries like rate/irate will struggle to process that many sequences.
  • Identify top label combinations driving cardinality
    Use topk to pinpoint which label pairs/groups are generating the most time series:
    topk(10, count by (label1, label2, label3) ({__name__="your_target_metric"}))
    
    Replace label1/2/3 with the actual labels in your metric. This will show you exactly which dimensions are exploding in size.
  • Check Prometheus TSDB status
    If you have access to the Prometheus UI, navigate to the /status page and look at the Time Series count. Alternatively, use this metric to monitor real-time series count:
    prometheus_tsdb_head_series
    
    Compare the total series count to your metric’s contribution—if your metric makes up a large chunk, that’s a clear red flag.
  • Test incremental query simplification
    Start with a stripped-down query like rate(your_target_metric[1m]) (no labels). If this runs fast, add one label at a time and re-run. The moment the query slows down or times out, you’ve found the label dimension causing the problem.
二、Query Performance Analysis Tricks

Beyond cardinality checks, these techniques will help you diagnose and optimize slow queries:

  • Enable Prometheus query logging
    In your Prometheus config, set query_log_file to a path (e.g., /var/log/prometheus/query_logs.json). The log will record every query’s duration, number of series scanned, total samples processed, and more. This is the most precise way to see exactly what’s making your query slow.
  • Use explain to inspect query execution plans
    Prefix your query with explain to see how Prometheus processes it:
    explain rate(your_target_metric{label="value"}[1m])
    
    The output will show if Prometheus is filtering labels before calculating rate (optimal) or vice versa (inefficient). This helps you spot unoptimized query logic.
  • Monitor Prometheus’s own performance metrics
    Keep an eye on these key metrics to correlate query behavior with TSDB health:
    • prometheus_engine_query_duration_seconds:Tracks query latency percentiles—see where your query falls (e.g., p99 over 60s confirms timeouts).
    • prometheus_tsdb_query_scrapes_total:Counts total samples scanned per query. Higher numbers directly correlate with slower queries.
    • prometheus_tsdb_head_samples_appended_total:Measures incoming sample volume—spikes here can overload the TSDB and slow down queries.
  • Use Grafana’s built-in query inspection
    In Grafana’s query editor, click the "Inspect" button (top-right), then select the "Query" tab. You’ll see the raw PromQL request, response time, and number of time series returned. If the series count is in the thousands, that’s the bottleneck.
  • Optimize query order: Aggregate first, calculate rate later
    Instead of rate(your_metric[1m]) by (label), try reversing the order:
    sum by (label) (rate(your_metric[1m]))
    
    Aggregating reduces the number of time series before applying rate, which drastically cuts down processing time.
  • Tweak Prometheus TSDB settings (last resort)
    If cardinality is unavoidable, adjust storage configs like storage.tsdb.retention.time to shorten data retention, or storage.tsdb.max-block-duration to optimize block storage. Only do this after optimizing queries and labels, as it affects all metrics.

内容的提问来源于stack exchange,提问作者vrtx54234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:23:21