GKE Autopilot无法调度GPU Pod问题求助
GKE Autopilot集群GPU PodPending问题排查方案
问题背景
现有europe-west1区域的GKE Autopilot集群,运行Node.js和Python应用。为Python应用添加AI模型处理需求后,更新Deployment配置请求NVIDIA T4 GPU,但Pod始终处于Pending状态。
原Deployment配置
apiVersion: apps/v1 kind: Deployment metadata: name: <name> spec: replicas: 1 selector: matchLabels: app: <app> template: metadata: labels: app: <app> spec: serviceAccountName: <gke-sa> nodeSelector: cloud.google.com/gke-accelerator: "nvidia-tesla-t4" cloud.google.com/gke-accelerator-count: "1" containers: - name: gpu-pipeline-prod image: <docker image> imagePullPolicy: Always ports: - containerPort: 8080 env: - name: PORT value: "8080" ... resources: limits: nvidia.com/gpu: 1 requests: memory: "18Gi" cpu: "18" ephemeral-storage: "32Gi"
Pod Describe关键事件输出
Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal LoadBalancerNegNotReady 19s neg-readiness-reflector Waiting for pod to become healthy in at least one of the NEG(s): [<NEG name>] Warning FailedScheduling 16s (x2 over 19s) gke.io/optimize-utilization-scheduler 0/2 nodes are available: 2 node(s) didn't match Pod's node affinity/selector. preemption: 0/2 nodes are available: 2 Preemption is not helpful for scheduling.. Normal NotTriggerScaleUp 16s cluster-autoscaler pod didn't trigger scale-up (it wouldn't fit if a new node is added): 18 node(s) didn't match Pod's node affinity/selector, 1 node(s) had untolerated taint {cloud.google.com/gke-quick-remove: true}
问题分析
- 节点选择器配置无效:
cloud.google.com/gke-accelerator-count并非GKE Autopilot支持的标准节点选择器标签,会导致调度器无法匹配节点。 - 资源请求超出单节点规格:原配置中请求的18vCPU远超过NVIDIA T4对应的Autopilot节点(如n1-standard-4仅4vCPU),即使Autopilot自动调整了资源请求,仍存在规格不匹配问题。
- 区域GPU资源限制:europe-west1区域可能存在NVIDIA T4库存不足的情况,导致无法扩容出符合要求的节点。
- 亲和性与容忍度限制:Pod的
tolerations中包含cloud.google.com/machine-family:NoSchedule,可能限制了可选的机器家族范围。
解决步骤
- 修正节点选择器:移除
cloud.google.com/gke-accelerator-count: "1",仅保留cloud.google.com/gke-accelerator: "nvidia-tesla-t4",Autopilot会自动分配带有1块T4 GPU的节点。 - 调整资源请求至合理范围:根据T4对应的Autopilot节点规格(如n1-standard-4:4vCPU、15Gi内存),将Pod的CPU和内存请求调整至节点可承载范围内,示例配置:
注意:Autopilot要求GPU Pod的资源请求与limits一致,确保QoS为Guaranteed。resources: limits: nvidia.com/gpu: 1 cpu: "4" memory: "14Gi" requests: nvidia.com/gpu: 1 cpu: "4" memory: "14Gi" ephemeral-storage: "32Gi" - 检查区域GPU可用性:执行命令确认europe-west1各可用区的T4库存:
gcloud compute accelerator-types list --filter="zone:europe-west1* AND name=nvidia-tesla-t4" - 优化容忍度配置:若无需限制机器家族,移除
tolerations中的cloud.google.com/machine-family:NoSchedule规则,扩大节点选择范围。 - 验证集群GPU支持:检查集群是否允许创建GPU节点:
若未启用GPU支持,可通过命令更新集群配置(Autopilot支持动态启用GPU)。gcloud container clusters describe <cluster-name> --region europe-west1 | grep -i gpu
内容的提问来源于stack exchange,提问作者Loïc Combis
相关产品推荐
相关产品推荐

