Prometheus规则需求:容器CPU/内存持续增长时触发告警
实现容器CPU/内存持续增长的Prometheus告警规则(单表达式)
要实现容器CPU/内存持续增长时触发告警、波动时不触发的需求,可以通过单个PromQL表达式结合delta、count_over_time等函数完成,核心逻辑是验证指定窗口内所有相邻样本的增量均非负,且整体有实际增长。
完整告警表达式
# 内存持续增长逻辑 (sum by (container_id, pod) (container_memory_working_set_bytes) and sum by (container_id, pod) ( count_over_time(delta(container_memory_working_set_bytes[1m]) >= 0 [1h]) == count_over_time(container_memory_working_set_bytes[1h]) - 1 ) and sum by (container_id, pod) (delta(container_memory_working_set_bytes[1h]) > 1048576)) # 或CPU持续增长逻辑 or (sum by (container_id, pod) (rate(container_cpu_usage_seconds_total[1m])) and sum by (container_id, pod) ( count_over_time(delta(rate(container_cpu_usage_seconds_total[1m])[1m:1m]) >= 0 [1h]) == count_over_time(rate(container_cpu_usage_seconds_total[1m])[1h]) - 1 ) and sum by (container_id, pod) (delta(rate(container_cpu_usage_seconds_total[1m])[1h]) > 0.01))
表达式拆解说明
通用逻辑(CPU/内存共用)
- 按容器聚合指标:
sum by (container_id, pod)确保按单个容器维度监控,避免同一Pod内多个容器的指标干扰。 - 验证无下降趋势:
delta(...[1m]) >= 0:计算相邻1分钟样本的增量,判断是否非负(无下降)。count_over_time(...) == count_over_time(...) -1:统计窗口内满足“无下降”的样本对数,等于总相邻样本对数(总样本数-1),说明整个窗口内所有相邻节点都没有下降。
- 排除无增长的持平情况:最后一个
delta(...[1h]) > 阈值确保窗口内指标有实际增长,避免因指标一直持平误触发告警。
内存专用部分
- 用
container_memory_working_set_bytes(容器实际使用的内存,含缓存)作为监控指标,阈值设为1048576(1MB),可根据业务调整。
CPU专用部分
- 先通过
rate(container_cpu_usage_seconds_total[1m])计算每秒CPU使用率(计数器转速率),再对这个速率做增量判断;阈值0.01代表1小时内CPU使用率增长超过0.01核,可按需调整。
告警规则配置示例
将上述表达式放入Prometheus告警规则文件:
groups: - name: container_resource_leak rules: - alert: ContainerResourceContinuousGrowth expr: | (sum by (container_id, pod) (container_memory_working_set_bytes) and sum by (container_id, pod) ( count_over_time(delta(container_memory_working_set_bytes[1m]) >= 0 [1h]) == count_over_time(container_memory_working_set_bytes[1h]) - 1 ) and sum by (container_id, pod) (delta(container_memory_working_set_bytes[1h]) > 1048576)) or (sum by (container_id, pod) (rate(container_cpu_usage_seconds_total[1m])) and sum by (container_id, pod) ( count_over_time(delta(rate(container_cpu_usage_seconds_total[1m])[1m:1m]) >= 0 [1h]) == count_over_time(rate(container_cpu_usage_seconds_total[1m])[1h]) - 1 ) and sum by (container_id, pod) (delta(rate(container_cpu_usage_seconds_total[1m])[1h]) > 0.01)) for: 5m labels: severity: critical annotations: summary: "容器资源持续增长({{ $labels.pod }}/{{ $labels.container_id }})" description: "{{ $labels.pod }}中的容器{{ $labels.container_id }}在1小时内CPU/内存持续增长,无下降趋势,疑似内存泄漏。"
注意事项
- 时间窗口和阈值需根据业务场景调整:比如短周期监控可把
1h改成30m,阈值根据容器资源配额调整。 - 确保Prometheus的采样间隔(
scrape_interval)与表达式中的1m匹配,避免样本不足导致判断失效。
内容的提问来源于stack exchange,提问作者jan
相关产品推荐
相关产品推荐

