You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes Prometheus告警规则求助:容器内存超节点总容量

Fixing Prometheus Alert Rule for Container Memory vs Node Capacity

Hey there! Let's break down why your current PromQL expression isn't giving you the per-node results you need, and fix it so you can get accurate alerts when a node's container memory usage exceeds its total capacity.

The Problem With Your Current Expression

Your existing query:

sum(container_memory_usage_bytes{instance=~"sa.*.domain"}) >= sum(kube_node_status_capacity_memory_bytes{node=~"sa.*.domain"})

This calculates the global total of container memory usage across all matching instances, then compares it to the global total of node memory capacity across all matching nodes. That's why you only get a single numeric result—you're comparing two aggregate sums, not evaluating each node individually.

This approach is flawed because:

  • It won't alert you if a single node's containers are using all its memory (but the global total is still under the global capacity)
  • It might trigger false alerts if the global totals are close, even though no individual node is over capacity

The Correct PromQL Expression

You need to aggregate container memory usage per node, then compare each node's total container usage to its own memory capacity. Here's the right query:

sum by (node) (container_memory_usage_bytes{instance=~"sa.*.domain"}) >= on(node) group_left() kube_node_status_capacity_memory_bytes{node=~"sa.*.domain"}

Let's break this down:

  • sum by (node) (container_memory_usage_bytes{...}): Sums up all container memory usage for each individual node, preserving the node label so we can match it to the node's capacity.
  • >= on(node) group_left(): Matches the summed container usage (left side) to the corresponding node's capacity (right side) using the node label. This ensures we're comparing each node's container usage to its own total memory, not a global sum.

For clarity in alerting (to get a 1/0 boolean result instead of raw byte values), you can add bool:

sum by (node) (container_memory_usage_bytes{instance=~"sa.*.domain"}) >= bool on(node) group_left() kube_node_status_capacity_memory_bytes{node=~"sa.*.domain"}

Example Alert Rule

Here's how to wrap this into a complete Prometheus alert rule:

groups:
- name: node-resource-alerts
  rules:
  - alert: NodeContainerMemoryExceedsCapacity
    expr: sum by (node) (container_memory_usage_bytes{instance=~"sa.*.domain"}) >= bool on(node) group_left() kube_node_status_capacity_memory_bytes{node=~"sa.*.domain"}
    for: 2m
    labels:
      severity: critical
    annotations:
      summary: "Node {{ $labels.node }}: Container memory exceeds node capacity"
      description: "Total container memory usage on node {{ $labels.node }} is {{ humanize $value }} bytes, which meets or exceeds the node's total memory capacity of {{ humanize (kube_node_status_capacity_memory_bytes{node=$labels.node}) }} bytes."

This rule will trigger a critical alert for each individual node where containers are using all available memory, after the condition persists for 2 minutes (to avoid flapping).

内容的提问来源于stack exchange,提问作者Ronny Forberger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 11:57:47