You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GKE Autopilot StatefulSet扩容失败:Pod调度异常求助

GKE Autopilot StatefulSet 第三个Pod无法调度(Pending状态)

问题现象

在GKE Autopilot集群部署3副本Nginx StatefulSet后,仅2个副本正常运行,第三个Pod持续处于Pending状态,无法完成调度。

核心报错

调度失败提示:

FailedScheduling  77s (x3 over 11m)  gke.io/optimize-utilization-scheduler  0/4 nodes are available: 4 Insufficient cpu, 4 Insufficient memory. preemption: 0/4 nodes are available: 4 No preemption victims found for incoming pod.

节点扩容失败提示:

FailedScaleUp     4m32s              cluster-autoscaler                     Node scale up in zones us-central1-b associated with this pod failed: IP space exhausted. Pod is at risk of not being scheduled.

当前Pod状态

kubectl get pods                        
NAME                  READY   STATUS    RESTARTS   AGE
nginx-statefulset-0   2/2     Running   0          25m
nginx-statefulset-1   2/2     Running   0          24m
nginx-statefulset-2   0/2     Pending   0          10m

StatefulSet配置

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: nginx-statefulset
  namespace: default
  labels:
    app: nginx
spec:
  serviceName: "nginx"
  replicas: 3
  selector:
    matchLabels:
      app: nginx
  template:
    metadata:
      labels:
        app: nginx
    spec:
      containers:
        - name: nginx
          image: nginx:1.21
          ports:
            - containerPort: 80
              name: web
          volumeMounts:
            - name: www
              mountPath: /usr/share/nginx/html
  volumeClaimTemplates:
    - metadata:
        name: www
      spec:
        accessModes: ["ReadWriteOnce"]
        resources:
          requests:
            storage: 1Gi

已确认IAM配额使用率低于1%,但Autopilot未完成自动扩容。


问题根因

  1. VPC子网IP耗尽:集群所在的us-central1-b、us-central1-c可用区的子网IP地址已分配完毕,导致Autopilot无法创建新节点。
  2. 现有节点资源饱和:当前4个节点的CPU和内存已被占满,单个Pod(含Istio Sidecar)总CPU请求为850m、内存请求为2.25Gi,现有节点无剩余资源容纳该Pod。

解决方案

1. 扩展VPC子网IP范围

  • 检查集群关联的VPC子网配置,确认目标可用区的IP地址是否耗尽。
  • 若子网IP不足,直接扩展子网CIDR范围,或在未使用的可用区(如us-central1-f)添加新子网,为Autopilot提供新的节点创建空间。

2. 调整Pod资源请求(临时缓解)

当前Nginx容器的资源请求较高,可临时降低请求值,让现有节点能容纳第三个Pod:
修改StatefulSet模板,添加资源请求/限制配置:

containers:
  - name: nginx
    image: nginx:1.21
    ports:
      - containerPort: 80
        name: web
    resources:
      requests:
        cpu: "300m"
        memory: "1Gi"
      limits:
        cpu: "650m"
        memory: "2Gi"
    volumeMounts:
      - name: www
        mountPath: /usr/share/nginx/html

应用修改:

kubectl apply -f <your-statefulset-file.yaml>

3. 扩展集群可用区

确认集群是否启用了多可用区部署,若仅依赖少数可用区,容易出现IP耗尽问题。添加更多可用区到集群,让Autopilot可以在其他区域调度新节点。

4. 清理闲置资源

检查集群内是否存在闲置Pod、节点或其他资源,释放冗余资源后重新尝试调度。


内容的提问来源于stack exchange,提问作者Aparna Raman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 15:09:52