You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Prometheus-adapter自定义server_cpu指标未显示,EKS HPA配置受阻

解决EKS中Prometheus-adapter自定义指标server_cpu不显示问题

先确认Prometheus本身能拿到目标指标

先到Prometheus UI执行你用来计算5分钟CPU利用率的PromQL,比如:

avg(irate(container_cpu_usage_seconds_total{namespace="你的命名空间", pod=~"你的服务Pod前缀.*"}[5m])) by (pod) * 100

如果这个查询返回空,问题出在Prometheus指标采集环节——检查Pod的prometheus.io/scrape注解是否开启,kubelet或cadvisor是否正常采集容器指标,或者命名空间、Pod标签是否匹配。

检查Prometheus-adapter的ConfigMap配置

以下是适配5分钟CPU利用率的正确配置模板,对照你的配置找问题:

apiVersion: v1
kind: ConfigMap
metadata:
  name: prometheus-adapter
  namespace: monitoring
data:
  config.yaml: |-
    rules:
    - seriesQuery: 'container_cpu_usage_seconds_total{namespace!="",pod!=""}'
      resources:
        overrides:
          namespace: {resource: "namespace"}
          pod: {resource: "pod"}
      name:
        matches: "^container_cpu_usage_seconds_total$"
        as: "server_cpu"
      metricsQuery: 'avg(irate(<<.Series>>[5m])) by (<<.GroupBy>>) * 100'
  • seriesQuery必须匹配Prometheus中存在的原始指标,不能写错标签或指标名
  • metricsQuery里的<<.Series>>会自动替换为seriesQuery匹配到的指标,<<.GroupBy>>对应resources里的字段,确保这个PromQL在Prometheus里能返回结果
  • 确认name.as是server_cpu,拼写不能出错

查看Prometheus-adapter日志定位错误

重启adapter后实时查看日志,抓取配置或查询类错误:

kubectl logs -n monitoring deployment/prometheus-adapter -f

常见问题:

  • YAML配置缩进错误,导致解析失败
  • adapter的Prometheus地址配置错误,无法连接Prometheus(检查ConfigMap里的prometheus.url)
  • PromQL语法错误,返回空结果,adapter不会生成对应指标

验证adapter内部是否生成了指标

端口转发adapter的8080端口,直接查询它的metrics端点:

kubectl port-forward -n monitoring deployment/prometheus-adapter 8080:8080
curl http://localhost:8080/metrics | grep server_cpu

如果这里能看到server_cpu相关指标,说明adapter内部逻辑正常,下一步检查RBAC:确保adapter的ServiceAccount拥有custommetrics.metrics.k8s.io的读写权限,集群角色绑定配置正确。

指标正常后的HPA配置示例

等server_cpu出现在kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1"结果里,用以下HPA配置实现5分钟CPU利用率超70%时的扩缩容:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: your-service-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: your-service
  minReplicas: 2
  maxReplicas: 10
  metrics:
  - type: Pods
    pods:
      metric:
        name: server_cpu
      target:
        type: AverageValue
        averageValue: 70

注意使用autoscaling/v2版本,v1版本不支持自定义指标。

EKS专属排查点

  • 确认EKS集群版本≥1.18,旧版本对自定义指标的支持有限
  • 若使用IAM Roles for Service Accounts(IRSA),确保adapter的SA有访问Prometheus的权限(如果Prometheus也用IRSA)
  • 检查VPC网络配置:adapter和Prometheus需在同一个VPC,或通过安全组允许两者通信

内容的提问来源于stack exchange,提问作者Roman Dayneko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 04:43:19