You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Prometheus规则需求:容器CPU/内存持续增长时触发告警

实现容器CPU/内存持续增长的Prometheus告警规则(单表达式)

要实现容器CPU/内存持续增长时触发告警、波动时不触发的需求,可以通过单个PromQL表达式结合delta、count_over_time等函数完成,核心逻辑是验证指定窗口内所有相邻样本的增量均非负,且整体有实际增长。

完整告警表达式

# 内存持续增长逻辑
(sum by (container_id, pod) (container_memory_working_set_bytes)
 and sum by (container_id, pod) (
   count_over_time(delta(container_memory_working_set_bytes[1m]) >= 0 [1h])
   == count_over_time(container_memory_working_set_bytes[1h]) - 1
 )
 and sum by (container_id, pod) (delta(container_memory_working_set_bytes[1h]) > 1048576))
# 或CPU持续增长逻辑
or
(sum by (container_id, pod) (rate(container_cpu_usage_seconds_total[1m]))
 and sum by (container_id, pod) (
   count_over_time(delta(rate(container_cpu_usage_seconds_total[1m])[1m:1m]) >= 0 [1h])
   == count_over_time(rate(container_cpu_usage_seconds_total[1m])[1h]) - 1
 )
 and sum by (container_id, pod) (delta(rate(container_cpu_usage_seconds_total[1m])[1h]) > 0.01))

表达式拆解说明

通用逻辑(CPU/内存共用)

  1. 按容器聚合指标:sum by (container_id, pod)确保按单个容器维度监控,避免同一Pod内多个容器的指标干扰。
  2. 验证无下降趋势:
    • delta(...[1m]) >= 0:计算相邻1分钟样本的增量,判断是否非负(无下降)。
    • count_over_time(...) == count_over_time(...) -1:统计窗口内满足“无下降”的样本对数,等于总相邻样本对数(总样本数-1),说明整个窗口内所有相邻节点都没有下降。
  3. 排除无增长的持平情况:最后一个delta(...[1h]) > 阈值确保窗口内指标有实际增长,避免因指标一直持平误触发告警。

内存专用部分

  • 用container_memory_working_set_bytes(容器实际使用的内存,含缓存)作为监控指标,阈值设为1048576(1MB),可根据业务调整。

CPU专用部分

  • 先通过rate(container_cpu_usage_seconds_total[1m])计算每秒CPU使用率(计数器转速率),再对这个速率做增量判断;阈值0.01代表1小时内CPU使用率增长超过0.01核,可按需调整。

告警规则配置示例

将上述表达式放入Prometheus告警规则文件:

groups:
- name: container_resource_leak
  rules:
  - alert: ContainerResourceContinuousGrowth
    expr: |
      (sum by (container_id, pod) (container_memory_working_set_bytes)
       and sum by (container_id, pod) (
         count_over_time(delta(container_memory_working_set_bytes[1m]) >= 0 [1h])
         == count_over_time(container_memory_working_set_bytes[1h]) - 1
       )
       and sum by (container_id, pod) (delta(container_memory_working_set_bytes[1h]) > 1048576))
      or
      (sum by (container_id, pod) (rate(container_cpu_usage_seconds_total[1m]))
       and sum by (container_id, pod) (
         count_over_time(delta(rate(container_cpu_usage_seconds_total[1m])[1m:1m]) >= 0 [1h])
         == count_over_time(rate(container_cpu_usage_seconds_total[1m])[1h]) - 1
       )
       and sum by (container_id, pod) (delta(rate(container_cpu_usage_seconds_total[1m])[1h]) > 0.01))
    for: 5m
    labels:
      severity: critical
    annotations:
      summary: "容器资源持续增长({{ $labels.pod }}/{{ $labels.container_id }})"
      description: "{{ $labels.pod }}中的容器{{ $labels.container_id }}在1小时内CPU/内存持续增长,无下降趋势,疑似内存泄漏。"

注意事项

  • 时间窗口和阈值需根据业务场景调整:比如短周期监控可把1h改成30m,阈值根据容器资源配额调整。
  • 确保Prometheus的采样间隔(scrape_interval)与表达式中的1m匹配,避免样本不足导致判断失效。

内容的提问来源于stack exchange,提问作者jan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 08:40:29