You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keda ScaledObject指标与Prometheus不符,如何实现GPU自动扩缩容?

问题

在云GPU服务商环境中创建Keda ScaledObject,环境通过Prometheus暴露指标,初始配置如下:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: [name]
  namespace: [namespace]
spec:
  cooldownPeriod: 30
  fallback:
    failureThreshold: 20
    replicas: 0
  maxReplicaCount: 4
  minReplicaCount: 1
  pollingInterval: 15
  scaleTargetRef:
    name: [deployment]
  triggers:
  - metadata:
      metricName: gpu-util
      metricType: Value
      query: |-
        avg(avg_over_time(DCGM_FI_DEV_GPU_UTIL[1m]))
      serverAddress: [address]:9090
      threshold: '80'
    type: prometheus

DCGM_FI_DEV_GPU_UTIL为NVIDIA GPU利用率指标,ScaledObject状态显示正常:

$ kubectl describe scaledobject [name] -n [namespace]
Name:         [name]
Namespace:    [namespace]
Labels:       scaledobject.keda.sh/name=[name]
Annotations:  <none>
API Version:  keda.sh/v1alpha1
Kind:         ScaledObject
Metadata:
  Creation Timestamp:  2023-04-28T01:27:50Z
  Finalizers:
    finalizer.keda.sh
  Generation:        1
  Resource Version:  36215438066
  UID:               [uid]
Spec:
  Cooldown Period:  30
  Fallback:
    Failure Threshold:  20
    Replicas:           0
  Max Replica Count:    4
  Min Replica Count:    1
  Polling Interval:     15
  Scale Target Ref:
    Name:  hashtop-1
  Triggers:
    Metadata:
      Metric Name:     gpu-util
      Namespace:       [namespace]
      Query:           avg(avg_over_time(DCGM_FI_DEV_GPU_UTIL[1m]))
      Server Address:  [url]:9090
      Threshold:       80
    Type:              prometheus
Status:
  Conditions:
    Message:  ScaledObject is defined correctly and is ready for scaling
    Reason:   ScaledObjectReady
    Status:   True
    Type:     Ready
    Message:  Scaling is not performed because triggers are not active
    Reason:   ScalerNotActive
    Status:   False
    Type:     Active
    Message:  No fallbacks are active on this scaled object
    Reason:   NoFallbackFound
    Status:   False
    Type:     Fallback
  External Metric Names:
    s0-prometheus-gpu-util
  Health:
    s0-prometheus-gpu-util:
      Number Of Failures:  0
      Status:              Happy
  Original Replica Count:  1
  Scale Target GVKR:
    Group:            apps
    Kind:             Deployment
    Resource:         deployments
    Version:          v1
  Scale Target Kind:  apps/v1.Deployment
Events:
  Type    Reason              Age                From           Message
  ----    ------              ----               ----           -------
  Normal  KEDAScalersStarted  78s                keda-operator  Started scalers watch
  Normal  ScaledObjectReady   63s (x2 over 78s)  keda-operator  ScaledObject is ready for scaling

直接在Prometheus执行查询可得到预期结果:GPU高负载时返回99,空闲时返回0:

# 高负载场景
$ curl '[url]/api/v1/query?query=avg(avg_over_time(DCGM_FI_DEV_GPU_UTIL\[1m\]))' | jq '.'
{
  "status": "success",
  "data": {
    "resultType": "vector",
    "result": [
      {
        "metric": {},
        "value": [
          1682728336,
          "99"
        ]
      }
    ]
  }
}

但Keda生成的HorizontalPodAutoscaler(HPA)显示的指标数据与Prometheus不符:GPU空闲时HPA显示18-20,4个GPU满负载时HPA显示36117500m,导致自动扩缩容逻辑失效,且无法直接访问Keda Operator。

需要修改ScaledObject的哪些配置,才能让HPA基于Prometheus的GPU利用率指标正常扩缩容?

解决方案

需调整以下ScaledObject配置项:

  • 修正Prometheus查询语句,关联目标Pod
    当前查询是全局GPU利用率平均值,未关联到目标Deployment的Pod,KEDA无法对应到具体扩缩容对象。修改查询语句,过滤出目标Deployment下的Pod指标:

    query: |-
      avg(avg_over_time(DCGM_FI_DEV_GPU_UTIL{pod=~"[deployment-name]-.*"}[1m])) by (pod)
    

    替换[deployment-name]为你的Deployment前缀,确保仅查询目标Pod的GPU数据。

  • 将metricType改为Utilization
    当前使用的Value类型基于指标绝对值扩缩容,GPU利用率属于资源利用率类指标,需用Utilization类型让KEDA基于百分比计算扩缩容比例:

    metricType: Utilization
    
  • 添加value参数(适配Utilization类型)
    使用Utilization类型时,需明确value参数指定目标利用率阈值,与threshold配合生效:

    triggers:
    - metadata:
        ...
        threshold: '80'
        value: '80'
        ...
    
  • 可选:调整指标单位匹配HPA预期
    若HPA仍显示异常数值,可将百分比转换为0-1的小数格式,此时阈值同步调整:

    query: |-
      avg(avg_over_time(DCGM_FI_DEV_GPU_UTIL{pod=~"[deployment-name]-.*"}[1m])) by (pod) / 100
    

    对应阈值改为0.8,代表80%利用率。

修改后的完整ScaledObject示例:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: [name]
  namespace: [namespace]
spec:
  cooldownPeriod: 30
  fallback:
    failureThreshold: 20
    replicas: 0
  maxReplicaCount: 4
  minReplicaCount: 1
  pollingInterval: 15
  scaleTargetRef:
    name: [deployment]
  triggers:
  - metadata:
      metricName: gpu-util
      metricType: Utilization
      query: |-
        avg(avg_over_time(DCGM_FI_DEV_GPU_UTIL{pod=~"[deployment]-.*"}[1m])) by (pod)
      serverAddress: [address]:9090
      threshold: '80'
      value: '80'
    type: prometheus

内容的提问来源于stack exchange,提问作者Jeffrey Mixon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 07:53:22