如何配置Grafana告警:耗时超3s的/createmag请求占比超5%触发
告警配置思路与解决方法
核心逻辑
要实现**15分钟内,/createmag接口POST 200请求中耗时超过3秒的请求占比>5%**的告警,核心是利用Prometheus直方图的分桶指标计算占比,步骤如下:
正确的PromQL语句
# 计算耗时>3秒的请求占比并判断阈值 ( increase(http_request_duration_seconds_count{handler="/createmag",method="POST",status="200"}[15m]) - increase(http_request_duration_seconds_bucket{handler="/createmag",method="POST",status="200",le="3.0"}[15m]) ) / increase(http_request_duration_seconds_count{handler="/createmag",method="POST",status="200"}[15m]) > 0.05
语句拆解
increase(http_request_duration_seconds_count[...]:获取15分钟内该接口的总请求数增量(count是累计计数器,increase用于计算时间窗口内的增量)increase(http_request_duration_seconds_bucket{le="3.0"}[...]:获取15分钟内耗时≤3秒的请求数增量(直方图的bucket是累计分桶,le="3.0"包含所有耗时≤3秒的请求)- 两者差值即为耗时>3秒的请求数,除以总请求数得到占比,最后判断是否超过5%(0.05)
你原有写法的问题
- 仅计算了耗时超3秒的请求数量,未计算占比,无法满足“比例超过5%”的告警条件
- 时间窗口使用了5m,不符合需求的15分钟统计范围
- 指标名简写为
http_xxx,未匹配实际的http_request_duration_seconds_前缀指标,导致数据不准确
额外注意事项
- 避免除以0:如果总请求数为0,会出现NaN,可通过
clamp_min处理,或添加总请求数阈值条件过滤小流量场景:# 增加总请求数≥100的条件,避免小流量下误告警 ( increase(http_request_duration_seconds_count{handler="/createmag",method="POST",status="200"}[15m]) - increase(http_request_duration_seconds_bucket{handler="/createmag",method="POST",status="200",le="3.0"}[15m]) ) / increase(http_request_duration_seconds_count{handler="/createmag",method="POST",status="200"}[15m]) > 0.05 and increase(http_request_duration_seconds_count{handler="/createmag",method="POST",status="200"}[15m]) > 100 - 验证指标存在:先在Prometheus查询界面确认
http_request_duration_seconds_bucket指标存在且带有正确的handler/method/status标签 - 测试语句有效性:等待有请求数据时,运行PromQL查看计算出的占比是否符合实际情况
内容的提问来源于stack exchange,提问作者SanchesAngryReut
相关产品推荐
相关产品推荐

