You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

HPA对接Prometheus-adapter获取自定义指标报无指标返回问题排查

HPA搭配Prometheus-adapter实现Kubernetes Pod弹性伸缩故障排查

故障背景

计划通过推送应用自定义扩缩容指标,基于该指标完成Pod水平自动扩缩,部署配置后出现HPA运行异常。

已观测到的现象

  • 可通过custom metrics API查询到自定义指标,但指标列表存在间歇性展示问题,查询命令与返回结果如下:
$ kubectl get --raw /apis/custom.metrics.k8s.io/v1beta1|jq 
{
  "kind": "APIResourceList",
  "apiVersion": "v1",
  "groupVersion": "custom.metrics.k8s.io/v1beta1",
  "resources": [
    {
      "name": "namespaces/http_requests_per_second",
      "singularName": "",
      "namespaced": false,
      "kind": "MetricValueList",
      "verbs": [
        "get"
      ]
    },
    {
      "name": "pods/http_requests_per_second",
      "singularName": "",
      "namespaced": true,
      "kind": "MetricValueList",
      "verbs": [
        "get"
      ]
    },
    {
      "name": "jobs.batch/http_requests_per_second",
      "singularName": "",
      "namespaced": true,
      "kind": "MetricValueList",
      "verbs": [
        "get"
      ]
    }
  ]
}
  • HPA持续抛出拉取指标失败告警,告警信息如下:
Warning  FailedGetPodsMetric  4m57s (x128 over 69m)  horizontal-pod-autoscaler  unable to get metric http_requests_per_second: no metrics returned from custom metrics API
  • 当前使用的HPA配置如下:
---
kind: HorizontalPodAutoscaler
apiVersion: autoscaling/v2beta2
metadata:
  name: app1
  namespace: default
  labels:
    app.kubernetes.io/name: app1
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: app1
  minReplicas: 1
  maxReplicas: 10
  metrics:
  - type: Pods
    pods:
      metric:
        name: http_requests_per_second
      target:
        type: AverageValue
        averageValue: 10m
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 80
  • 已将prometheus-adapter日志级别调整为7级排查,未发现错误、警告或可疑事件记录。

可能的故障诱因

  • 指标标签关联缺失:Prometheus存储的http_requests_per_second指标未携带pod标签,或pod标签值与Kubernetes集群中实际运行的Pod名称不匹配。Prometheus-adapter依赖pod、namespace标签将指标序列和对应工作负载的Pod做关联,标签缺失或不匹配时,即使API资源列表能展示指标定义,查询具体Pod维度的指标值时也会返回空结果。
  • Prometheus-adapter规则配置错误:adapter的自定义指标匹配规则未正确配置Pod维度关联逻辑,比如seriesQuery筛选条件未匹配到目标Pod对应的指标序列,或resources配置段未正确建立标签和Kubernetes Pod资源的映射关系,会导致指标列表偶发可见但实际查值为空。
  • 指标采样断流:应用侧推送http_requests_per_second指标存在间歇性中断,Prometheus中该指标在查询时间点没有有效采样点,adapter查询时会返回空结果,同时会造成API资源列表偶发不展示对应指标的现象。
  • 查询范围不匹配:HPA配置的是Pods类型指标,会查询default命名空间下属于app1 Deployment的所有Pod对应指标,如果指标未携带namespace标签,或标签值不为default,adapter无法筛选出目标Pod对应的指标序列,就会返回无指标的错误。
  • 多副本配置不一致:如果集群中部署了多副本Prometheus-adapter,部分副本存在规则配置错误、与Prometheus连接异常的问题,HPA轮询请求打到不同副本时,就会出现指标列表间歇性可见、查值偶发失败的现象。
  • 单位配置冲突:HPA中配置的averageValue: 10m是千分单位计量值,如果http_requests_per_second本身是无单位的比率型指标,adapter做单位转换时可能出现匹配异常,返回空结果。

内容的提问来源于stack exchange,提问作者Vishal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 01:24:32