You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否基于GPU指标实现Seldon Deployment的自动扩缩容?

基于GPU指标实现Seldon Deployment自动扩缩容

可以基于GPU指标实现Seldon Deployment的自动扩缩容,但无法仅依赖默认的metric-server,需结合自定义指标采集与Kubernetes HPA(水平Pod自动扩缩器)完成,具体步骤如下:

1. 确保GPU指标可采集

在AWS EKS集群已部署Nvidia设备插件的基础上,需补充指标采集组件:

  • 部署Nvidia DCGM-Exporter:用于采集GPU使用率、显存占用等核心指标,并以Prometheus格式暴露
  • 配置Prometheus:确保其能抓取DCGM-Exporter暴露的指标端点

2. 部署自定义指标适配器

需将Prometheus中的GPU指标转换为Kubernetes可识别的自定义指标,推荐使用Prometheus Adapter:

  • 将Prometheus Adapter部署至集群
  • 配置Adapter映射规则,把dcgm_gpu_utilization、dcgm_gpu_memory_used等GPU指标,映射为Kubernetes可识别的自定义指标(如pods/gpu_utilization)

3. 为Seldon Deployment配置HPA

创建HPA资源,指定基于GPU自定义指标触发扩缩容:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: seldon-model-hpa
spec:
  scaleTargetRef:
    apiVersion: machinelearning.seldon.io/v1
    kind: SeldonDeployment
    name: <你的Seldon Deployment名称>
  minReplicas: 1
  maxReplicas: 5
  metrics:
  - type: Pods
    pods:
      metric:
        name: gpu_utilization
      target:
        type: AverageValue
        averageValue: 70%
  • 可根据业务需求调整averageValue阈值(如70%),达到阈值时触发扩缩容
  • 确保HPA能正确识别Seldon Deployment作为扩缩容目标

4. 验证扩缩容逻辑

  • 给模型服务施加GPU负载,观察GPU使用率是否达阈值
  • 查看HPA状态,确认是否触发副本数调整
  • 检查Seldon Deployment的副本数变化,验证扩缩容生效

内容的提问来源于stack exchange,提问作者Petar R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 17:57:03