You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何查询AWS EKS节点组最大节点数并在Grafana展示对比及告警?

AWS EKS节点组最大节点数与当前节点数对比查询方案

核心思路

EKS节点组底层对应AWS Auto Scaling Group(ASG),节点组的最大节点数就是ASG的MaxSize配置值,当前运行节点数可通过ASG的GroupInServiceInstances指标获取,需借助CloudWatch Exporter将AWS CloudWatch指标导入Prometheus。

步骤1:部署CloudWatch Exporter

用Helm部署CloudWatch Exporter,并配置指标抓取规则(以myregion区域、mygroup节点组为例):

  1. 创建config.yml配置文件:
region: myregion
metrics:
- namespace: AWS/AutoScaling
  dimensions:
  - name: AutoScalingGroupName
    value: eks-mygroup-.*  # 匹配节点组对应的ASG名称,EKS节点组ASG命名规则一般为eks-节点组名-随机后缀
  metrics:
  - name: GroupInServiceInstances
    statistic: Sum
    period: 60
  - name: MaxSize
    statistic: Average
    period: 60
  1. 通过Helm安装:
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm install cloudwatch-exporter prometheus-community/prometheus-cloudwatch-exporter -f config.yml

步骤2:Prometheus查询语句

当前节点数

sum(aws_ec2_autoscaling_group_in_service_instances{auto_scaling_group_name=~"eks-mygroup-.*"})

最大节点数

sum(aws_ec2_autoscaling_group_max_size{auto_scaling_group_name=~"eks-mygroup-.*"})

节点使用率(用于对比展示)

sum(aws_ec2_autoscaling_group_in_service_instances{auto_scaling_group_name=~"eks-mygroup-.*"}) / sum(aws_ec2_autoscaling_group_max_size{auto_scaling_group_name=~"eks-mygroup-.*"})

步骤3:告警规则配置

在PrometheusRule中添加告警规则(示例:使用率超过80%且持续5分钟触发告警):

groups:
- name: eks-nodegroup-alerts
  rules:
  - alert: EKSNodeGroupHighUtilization
    expr: sum(aws_ec2_autoscaling_group_in_service_instances{auto_scaling_group_name=~"eks-mygroup-.*"}) / sum(aws_ec2_autoscaling_group_max_size{auto_scaling_group_name=~"eks-mygroup-.*"}) > 0.8
    for: 5m
    labels:
      severity: warning
    annotations:
      summary: "EKS节点组 {{ $labels.auto_scaling_group_name }} 使用率过高"
      description: "当前节点数: {{ $value | round 0 }},最大节点数: {{ sum(aws_ec2_autoscaling_group_max_size{auto_scaling_group_name=~\"eks-mygroup-.*\"}) | round 0 }},使用率: {{ $value * 100 | round 1 }}%"

原up查询排查

若你原有的sum(up{instance=~".*.myregion.compute.internal", eks_amazonaws_com_nodegroup="mygroup"})无结果,可按以下排查:

  • 执行up{instance=~".*.myregion.compute.internal"}查看所有节点标签,确认是否存在eks_amazonaws_com_nodegroup标签,且值为mygroup
  • 检查Prometheus抓取kubelet的配置,确保节点的kubernetes.io/metadata标签被正确采集

内容的提问来源于stack exchange,提问作者maxisalamone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 23:15:33