You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GKE Autopilot无法调度GPU Pod问题求助

GKE Autopilot集群GPU PodPending问题排查方案

问题背景

现有europe-west1区域的GKE Autopilot集群,运行Node.js和Python应用。为Python应用添加AI模型处理需求后,更新Deployment配置请求NVIDIA T4 GPU,但Pod始终处于Pending状态。

原Deployment配置

apiVersion: apps/v1
kind: Deployment
metadata:
  name: <name>
spec:
  replicas: 1
  selector:
    matchLabels:
      app: <app>
  template:
    metadata:
      labels:
        app: <app>
    spec:
      serviceAccountName: <gke-sa>
      nodeSelector:
        cloud.google.com/gke-accelerator: "nvidia-tesla-t4"
        cloud.google.com/gke-accelerator-count: "1"
      containers:
      - name: gpu-pipeline-prod
        image: <docker image>
        imagePullPolicy: Always
        ports:
        - containerPort: 8080
        env:
          - name: PORT
            value: "8080"
          ...
        resources:
          limits:
            nvidia.com/gpu: 1
          requests:
            memory: "18Gi"
            cpu: "18"
            ephemeral-storage: "32Gi"

Pod Describe关键事件输出

Events:
  Type     Reason                   Age                From                                   Message
  ----     ------                   ----               ----                                   -------
  Normal   LoadBalancerNegNotReady  19s                neg-readiness-reflector                Waiting for pod to become healthy in at least one of the NEG(s): [<NEG name>]
  Warning  FailedScheduling         16s (x2 over 19s)  gke.io/optimize-utilization-scheduler  0/2 nodes are available: 2 node(s) didn't match Pod's node affinity/selector. preemption: 0/2 nodes are available: 2 Preemption is not helpful for scheduling..
  Normal   NotTriggerScaleUp        16s                cluster-autoscaler                     pod didn't trigger scale-up (it wouldn't fit if a new node is added): 18 node(s) didn't match Pod's node affinity/selector, 1 node(s) had untolerated taint {cloud.google.com/gke-quick-remove: true}

问题分析

  1. 节点选择器配置无效:cloud.google.com/gke-accelerator-count并非GKE Autopilot支持的标准节点选择器标签,会导致调度器无法匹配节点。
  2. 资源请求超出单节点规格:原配置中请求的18vCPU远超过NVIDIA T4对应的Autopilot节点(如n1-standard-4仅4vCPU),即使Autopilot自动调整了资源请求,仍存在规格不匹配问题。
  3. 区域GPU资源限制:europe-west1区域可能存在NVIDIA T4库存不足的情况,导致无法扩容出符合要求的节点。
  4. 亲和性与容忍度限制:Pod的tolerations中包含cloud.google.com/machine-family:NoSchedule,可能限制了可选的机器家族范围。

解决步骤

  • 修正节点选择器:移除cloud.google.com/gke-accelerator-count: "1",仅保留cloud.google.com/gke-accelerator: "nvidia-tesla-t4",Autopilot会自动分配带有1块T4 GPU的节点。
  • 调整资源请求至合理范围:根据T4对应的Autopilot节点规格(如n1-standard-4:4vCPU、15Gi内存),将Pod的CPU和内存请求调整至节点可承载范围内,示例配置:
    resources:
      limits:
        nvidia.com/gpu: 1
        cpu: "4"
        memory: "14Gi"
      requests:
        nvidia.com/gpu: 1
        cpu: "4"
        memory: "14Gi"
        ephemeral-storage: "32Gi"
    
    注意:Autopilot要求GPU Pod的资源请求与limits一致,确保QoS为Guaranteed。
  • 检查区域GPU可用性:执行命令确认europe-west1各可用区的T4库存:
    gcloud compute accelerator-types list --filter="zone:europe-west1* AND name=nvidia-tesla-t4"
    
  • 优化容忍度配置:若无需限制机器家族,移除tolerations中的cloud.google.com/machine-family:NoSchedule规则,扩大节点选择范围。
  • 验证集群GPU支持:检查集群是否允许创建GPU节点:
    gcloud container clusters describe <cluster-name> --region europe-west1 | grep -i gpu
    
    若未启用GPU支持,可通过命令更新集群配置(Autopilot支持动态启用GPU)。

内容的提问来源于stack exchange,提问作者Loïc Combis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 13:55:27