关于Prometheus告警规则的咨询:status为transit持续1小时的配置
关于Prometheus告警规则的分析与优化
你的方案问题
你当前的规则increase(flight_api_calls_to_mq_total{status="transit"}[1h]) > 50逻辑不匹配需求:
- 该规则判断的是1小时内
transit状态的调用增量超过50,但你的核心需求是**transit状态持续存在超过1小时**。 - 即使
transit状态已经持续1小时,但如果这段时间内调用量很少(比如增量只有10),该规则不会触发告警,不符合你的预期。
更优实现方案
从指标名称_total判断,它大概率是**计数器(Counter)**类型,以下分两种核心场景给出优化规则:
情况1:指标为计数器(Counter)
计数器仅递增,status="transit"的样本出现后,只要有过调用就会保留非零值。要判断该状态持续超过1小时,可使用:
flight_api_calls_to_mq_total{status="transit"} > 0 and flight_api_calls_to_mq_total{status="transit"} offset 1h > 0
- 逻辑:当前存在
transit状态的有效样本,且1小时前也存在该状态的样本,说明状态持续了至少1小时。
情况2:指标为仪表盘(Gauge)
若指标值可动态变化(比如会重置为0),要判断transit状态在1小时内持续存在,可使用:
count_over_time(flight_api_calls_to_mq_total{status="transit"}[1h]) >= (3600 / <你的采样间隔>)
- 替换
<你的采样间隔>为实际的Prometheus scrape interval(比如默认15s则填15),计算出1小时内的预期样本数。只要该时间段内的样本数达到预期,说明状态持续存在。
额外扩展建议
- 如果需要同时关注
transit状态的活跃性(确保有调用发生),可在规则中增加增量阈值:
(flight_api_calls_to_mq_total{status="transit"} > 0 and flight_api_calls_to_mq_total{status="transit"} offset 1h > 0) and increase(flight_api_calls_to_mq_total{status="transit"}[1h]) > 0
- 若想确保
transit状态是首次出现且持续1小时,可结合absent函数:
flight_api_calls_to_mq_total{status="transit"} > 0 and absent(flight_api_calls_to_mq_total{status="transit"} offset 2h) == 1
内容的提问来源于stack exchange,提问作者Rafa S
相关产品推荐
相关产品推荐

