如何编写Prometheus查询实现容器内存使用率超80%告警?
容器内存使用率超80%的Prometheus告警查询语句
背景说明
已为全命名空间下所有容器配置资源请求与限制,配置示例:
resources: requests: cpu: "53m" memory: "46Mi" limits: cpu: "340m" memory: "70Mi"
此前使用的PromQL查询未达到预期效果:
container_memory_working_set_bytes{name!~".*prometheus.*", image!="", container_name!="POD"} / container_spec_memory_limit_bytes{name!~".*prometheus.*", image!="", container_name!="POD"} * 100
正确的告警查询语句
以下PromQL可准确计算容器内存使用率,直接用于Grafana告警规则(当使用率超过80%时触发):
100 * (container_memory_working_set_bytes{name!~".*prometheus.*", image!="", container_name!="POD"} / on(namespace, pod, container) group_left() container_spec_memory_limit_bytes{name!~".*prometheus.*", image!="", container_name!="POD"}) > 80
核心优化点
- 新增
on(namespace, pod, container) group_left()标签关联规则,解决原查询中因标签不匹配导致的指标关联错误问题 - 直接嵌入
> 80的阈值判断,无需额外处理即可作为告警触发条件 - 保留原有的容器过滤逻辑,排除Prometheus自身容器、无镜像容器及POD基础容器
内容的提问来源于stack exchange,提问作者Yogesh Bhagwatkar
相关产品推荐
相关产品推荐

