GKE Autopilot集群中Job Pod无法挂载GCP Filestore PVC求助
GKE Autopilot Job Pod挂载Filestore PVC失败,持续处于ContainerCreating状态
在GKE Autopilot集群上运行一个parallelism=50的Kubernetes Job,该Job所需存储超过了Autopilot单节点最大临时存储(10Gi)。因需要为Pod提供ReadWriteMany权限的存储,选择使用GCP Filestore创建可挂载的PVC(虽然Filestore能提供小于1TiB的最小实例容量会更合适),但Job Pod始终处于ContainerCreating状态。查看事件日志,发现是MountVolume.MountDevice失败导致的问题:
Warning FailedScheduling 11m gke.io/optimize-utilization-scheduler 0/12 nodes are available: 11 Insufficient memory, 12 Insufficient cpu. preemption: 0/12 nodes are available: 12 No preemption victims found for incoming pod.. Normal TriggeredScaleUp 11m cluster-autoscaler pod triggered scale-up Normal Scheduled 6m39s gke.io/optimize-utilization-scheduler Successfully assigned default/mypod-7l5k9 to gk3-mycluster-3-e79620bd-jvsg Warning FailedMount 4m8s (x6 over 4m39s) kubelet MountVolume.MountDevice failed for volume "pvc-435bf565-25f0-43f7-86d4-b3ecadce43a3" : rpc error: code = Aborted desc = An operation with the given volume key modeInstance/asia-northeast1-b/pvc-435bf565-25f0-43f7-86d4-b3ecadce43a3/vol1 already exists. --- Most likely a long process is still running to completion. Retrying. Warning FailedMount 2m19s kubelet Unable to attach or mount volumes: unmounted volumes=[my-mounted-storage], unattached volumes=[kube-api-access-4gs6h shared-storage]: timed out waiting for the condition Warning FailedMount 96s (x2 over 4m39s) kubelet MountVolume.MountDevice failed for volume "pvc-435bf565-25f0-43f7-86d4-b3ecadce43a3" : rpc error: code = DeadlineExceeded desc = context deadline exceeded Warning FailedMount 5s (x2 over 4m36s) kubelet Unable to attach or mount volumes: unmounted volumes=[my-mounted-storage], unattached volumes=[my-mounted-storage kube-api-access-4gs6h]: timed out waiting for the condition
PVC清单
kind: PersistentVolumeClaim apiVersion: v1 metadata: name: podpvc spec: accessModes: - ReadWriteMany storageClassName: standard-rwx resources: requests: storage: 1Ti
Job清单
apiVersion: batch/v1 kind: Job metadata: name: mypod labels: app.kubernetes.io/name: mypod spec: parallelism: 50 template: metadata: name: mypod spec: serviceAccountName: workload-identity-sa volumes: - name: my-mounted-storage persistentVolumeClaim: claimName: podpvc containers: - name: mypod-container image: mypod-image:staging-0.1 imagePullPolicy: Always env: - name: env value: "stg" resources: requests: cpu: "4" memory: "16Gi" volumeMounts: - name: my-mounted-storage mountPath: /mnt/data restartPolicy: OnFailure
已进行的排查操作:
- PV和PVC状态均为健康且已绑定
- 节点无现存卷挂载(执行
kubectl describe nodes | grep Attach未找到相关内容) - 删除PVC和Job后重新创建,问题依旧存在

内容的提问来源于stack exchange,提问作者Coding Tumbleweed
相关产品推荐
相关产品推荐

