You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GKE集群如何实现Pod优先调度GPU节点无资源时调度CPU节点

问题根因

你当前的配置无法调度到CPU节点的核心原因是在容器资源限制中声明了nvidia.com/gpu: 1,该配置属于Kubernetes调度的硬约束,调度器只会匹配拥有对应GPU资源的节点,自然无法调度到无GPU的CPU节点。

推荐解决方案:双Deployment + 优先级类(适配GKE 1.19版本,生产可用)

核心逻辑是拆分部署为GPU优先版和CPU兜底版,通过优先级配置实现GPU资源充足时优先跑GPU Pod,GPU不足时自动调度CPU Pod。

步骤1:创建高优先级类(供GPU Pod使用)

优先级更高的GPU Pod可以在GPU节点释放资源时,优先抢占节点上运行的低优先级CPU Pod,确保GPU资源被充分利用:

apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: gpu-pod-priority
value: 100000
globalDefault: false
description: "GPU优先部署的Pod优先级"

步骤2:创建GPU版Deployment

保留你原有GPU相关配置,新增优先级字段,强制调度到GPU节点:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: NAME-gpu
  namespace: NAMESPACE
spec:
  replicas: 3 # 按业务需要调整副本数
  selector:
    matchLabels:
      app: NAME
      runtime: gpu
  template:
    metadata:
      labels:
        app: NAME
        runtime: gpu
    spec:
      priorityClassName: gpu-pod-priority
      nodeSelector:
        cloud.google.com/gke-preemptible: "true"
      affinity:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
            - matchExpressions:
              - key: cloud.google.com/gke-accelerator
                operator: In
                values:
                - nvidia-tesla-t4
      containers:
      - name: NAME
        image: IMAGE # 镜像必须同时兼容CPU/GPU运行
        resources:
          requests:
            memory: 28.0Gi
            cpu: 3000m
          limits:
            cpu: 4000m
            nvidia.com/gpu: 1
      tolerations:
      - effect: NoSchedule
        key: nvidia.com/gpu
        operator: Exists

步骤3:创建CPU兜底版Deployment

不配置GPU资源约束,优先调度到CPU节点,使用默认低优先级:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: NAME-cpu
  namespace: NAMESPACE
spec:
  replicas: 3 # 可与GPU版副本数一致,也可单独配置HPA自动扩缩
  selector:
    matchLabels:
      app: NAME
      runtime: cpu
  template:
    metadata:
      labels:
        app: NAME
        runtime: cpu
    spec:
      nodeSelector:
        cloud.google.com/gke-preemptible: "true"
      affinity:
        nodeAffinity:
          preferredDuringSchedulingIgnoredDuringExecution:
          - weight: 100
            preference:
              matchExpressions:
              - key: cloud.google.com/gke-accelerator
                operator: DoesNotExist
      containers:
      - name: NAME
        image: IMAGE # 和GPU版使用同一个兼容镜像
        resources:
          requests:
            memory: 28.0Gi
            cpu: 3000m
          limits:
            cpu: 4000m

可选优化:统一流量入口

如果需要对外提供统一服务,创建一个Service匹配app: NAME标签即可,自动将流量转发到所有正常运行的GPU/CPU Pod:

apiVersion: v1
kind: Service
metadata:
  name: NAME-svc
  namespace: NAMESPACE
spec:
  selector:
    app: NAME
  ports:
  - port: 80
    targetPort: 8080 # 按业务实际暴露端口调整

注意事项

  • 确保业务镜像同时兼容CPU和GPU运行环境,启动时可以自动检测GPU设备是否存在,自动切换对应运行模式
  • 抢占式节点被回收时,K8s会自动重新调度Pod,优先尝试调度GPU版,失败则调度CPU版,完全匹配你的需求
  • 如果不想维护两个Deployment,也可以自行部署动态准入控制器,在Pod创建时检测可用节点自动注入/取消GPU资源约束,但该方案复杂度更高,不推荐生产环境使用

内容的提问来源于stack exchange,提问作者Montoya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 09:36:01