You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GKE Autopilot集群中Job Pod无法挂载GCP Filestore PVC求助

GKE Autopilot Job Pod挂载Filestore PVC失败,持续处于ContainerCreating状态

在GKE Autopilot集群上运行一个parallelism=50的Kubernetes Job,该Job所需存储超过了Autopilot单节点最大临时存储(10Gi)。因需要为Pod提供ReadWriteMany权限的存储,选择使用GCP Filestore创建可挂载的PVC(虽然Filestore能提供小于1TiB的最小实例容量会更合适),但Job Pod始终处于ContainerCreating状态。查看事件日志,发现是MountVolume.MountDevice失败导致的问题:

Warning  FailedScheduling  11m                   gke.io/optimize-utilization-scheduler  0/12 nodes are available: 11 Insufficient memory, 12 Insufficient cpu. preemption: 0/12 nodes are available: 12 No preemption victims found for incoming pod..
Normal   TriggeredScaleUp  11m                   cluster-autoscaler                     pod triggered scale-up
Normal   Scheduled         6m39s                 gke.io/optimize-utilization-scheduler  Successfully assigned default/mypod-7l5k9 to gk3-mycluster-3-e79620bd-jvsg
Warning  FailedMount       4m8s (x6 over 4m39s)  kubelet                                MountVolume.MountDevice failed for volume "pvc-435bf565-25f0-43f7-86d4-b3ecadce43a3" : rpc error: code = Aborted desc = An operation with the given volume key modeInstance/asia-northeast1-b/pvc-435bf565-25f0-43f7-86d4-b3ecadce43a3/vol1 already exists.
--- Most likely a long process is still running to completion. Retrying.
Warning  FailedMount  2m19s                kubelet  Unable to attach or mount volumes: unmounted volumes=[my-mounted-storage], unattached volumes=[kube-api-access-4gs6h shared-storage]: timed out waiting for the condition
Warning  FailedMount  96s (x2 over 4m39s)  kubelet  MountVolume.MountDevice failed for volume "pvc-435bf565-25f0-43f7-86d4-b3ecadce43a3" : rpc error: code = DeadlineExceeded desc = context deadline exceeded
Warning  FailedMount  5s (x2 over 4m36s)   kubelet  Unable to attach or mount volumes: unmounted volumes=[my-mounted-storage], unattached volumes=[my-mounted-storage kube-api-access-4gs6h]: timed out waiting for the condition

PVC清单

kind: PersistentVolumeClaim
apiVersion: v1
metadata:
  name: podpvc
spec:
  accessModes:
    - ReadWriteMany
  storageClassName: standard-rwx
  resources:
    requests:
      storage: 1Ti

Job清单

apiVersion: batch/v1
kind: Job
metadata:
  name: mypod
  labels:
    app.kubernetes.io/name: mypod
spec:
  parallelism: 50
  template:
    metadata:
      name: mypod
    spec:
      serviceAccountName: workload-identity-sa
      volumes:
      - name: my-mounted-storage
        persistentVolumeClaim:
          claimName: podpvc
      containers:
      - name: mypod-container
        image: mypod-image:staging-0.1
        imagePullPolicy: Always
        env:
        - name: env
          value: "stg"
        resources:
          requests:
            cpu: "4"
            memory: "16Gi"
        volumeMounts:
        - name: my-mounted-storage
          mountPath: /mnt/data
      restartPolicy: OnFailure

已进行的排查操作:

  • PV和PVC状态均为健康且已绑定
  • 节点无现存卷挂载(执行kubectl describe nodes | grep Attach未找到相关内容)
  • 删除PVC和Job后重新创建,问题依旧存在

Filestore状态截图

内容的提问来源于stack exchange,提问作者Coding Tumbleweed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 13:35:38