You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GKE中Filestore CSI卷挂载Pod超时失败求助

卷挂载超时问题排查与解决

问题现象

Pod启动时出现卷挂载超时错误:

Unable to attach or mount volumes: unmounted volumes=[file-store], unattached volumes=[file-store kube-api-access-p7btw]: timed out waiting for the condition

排查信息

Pod描述信息

kubectl describe pod kickstar-backend-5577cf96cf-rlpht
Name:             kickstar-backend-5577cf96cf-rlpht
Namespace:        default
Priority:         0
Service Account:  default
Node:             gke-kickstar-prod-workloads-clus-main-ab940238-1ktc/10.128.0.5
Start Time:       Tue, 08 Aug 2023 03:31:56 +0000
Labels:           app=kickstar-backend
                  pod-template-hash=5577cf96cf
Annotations:      <none>
Status:           Pending
IP:                
IPs:              <none>
Controlled By:    ReplicaSet/kickstar-backend-5577cf96cf
Containers:
  kickstar-backend:
    Container ID:   
    Image:          asia-southeast1-docker.pkg.dev/kickstar-prod/kickstar/kickstar.backend:13fb26c
    Image ID:       
    Port:           1026/TCP
    Host Port:      0/TCP
    State:          Waiting
      Reason:       ContainerCreating
    Ready:          False
    Restart Count:  0
    Environment:    <none>
    Mounts:
      /app/wwwroot from file-store (rw)
      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-p7btw (ro)
Readiness Gates:
  Type                                       Status
  cloud.google.com/load-balancer-neg-ready   True 
Conditions:
  Type                                       Status
  cloud.google.com/load-balancer-neg-ready   True 
  Initialized                                True 
  Ready                                      False 
  ContainersReady                            False 
  PodScheduled                               True 
Volumes:
  file-store:
    Type:       PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
    ClaimName:  file-store-pvc
    ReadOnly:   false
  kube-api-access-p7btw:
    Type:                    Projected (a volume that contains injected data from multiple sources)
    TokenExpirationSeconds:  3607
    ConfigMapName:           kube-root-ca.crt
    ConfigMapOptional:       <nil>
    DownwardAPI:             true
QoS Class:                   BestEffort
Node-Selectors:              <none>
Tolerations:                 node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
                             node.kubernetes.io/unreachable:NoExecute op=Exists for 300s
Events:
  Type     Reason                   Age                From                     Message
  ----     ------                   ----               ----                     -------
  Normal   LoadBalancerNegNotReady  27m (x2 over 27m)  neg-readiness-reflector  Waiting for pod to become healthy in at least one of the NEG(s): [k8s1-0912e870-default-kickstar-backend-80-4d4b640a]
  Normal   NotTriggerScaleUp        27m                cluster-autoscaler       pod didn't trigger scale-up:
  Warning  FailedScheduling         24m (x2 over 27m)  default-scheduler        0/1 nodes are available: 1 pod has unbound immediate PersistentVolumeClaims. preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling.
  Normal   Scheduled                24m                default-scheduler        Successfully assigned default/kickstar-backend-5577cf96cf-rlpht to gke-kickstar-prod-workloads-clus-main-ab940238-1ktc
  Warning  FailedMount              21m (x6 over 22m)  kubelet                  MountVolume.MountDevice failed for volume "pvc-9cd5a129-d237-4074-95ae-de42446fc75c" : rpc error: code = Aborted desc = An operation with the given volume key modeInstance/asia-southeast1-a/pvc-9cd5a129-d237-4074-95ae-de42446fc75c/vol1 already exists.
 --- Most likely a long process is still running to completion. Retrying.
  Normal   LoadBalancerNegTimeout  12m                  neg-readiness-reflector  Timeout waiting for pod to become healthy in at least one of the NEG(s): [k8s1-0912e870-default-kickstar-backend-80-4d4b640a]. Marking condition "cloud.google.com/load-balancer-neg-ready" to True.
  Warning  FailedMount             4m12s (x2 over 11m)  kubelet                  Unable to attach or mount volumes: unmounted volumes=[file-store], unattached volumes=[kube-api-access-p7btw file-store]: timed out waiting for the condition
  Warning  FailedMount             118s (x8 over 22m)   kubelet                  Unable to attach or mount volumes: unmounted volumes=[file-store], unattached volumes=[file-store kube-api-access-p7btw]: timed out waiting for the condition
  Warning  FailedMount             5s (x7 over 22m)     kubelet                  MountVolume.MountDevice failed for volume "pvc-9cd5a129-d237-4074-95ae-de42446fc75c" : rpc error: code = DeadlineExceeded desc = context deadline exceeded

PV和PVC状态

kubectl get pv,pvc
NAME                                                        CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS   CLAIM                    STORAGECLASS   REASON   AGE
persistentvolume/pvc-9cd5a129-d237-4074-95ae-de42446fc75c   1Ti        RWX            Delete           Bound    default/file-store-pvc   file-store              31m

NAME                                   STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   AGE
persistentvolumeclaim/file-store-pvc   Bound    pvc-9cd5a129-d237-4074-95ae-de42446fc75c   1Ti        RWX            file-store     34m

StorageClass、PVC和Deployment配置

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: file-store
provisioner: filestore.csi.storage.gke.io
volumeBindingMode: Immediate
allowVolumeExpansion: true
parameters:
  tier: standard
  network: default
---
kind: PersistentVolumeClaim
apiVersion: v1
metadata:
  name: file-store-pvc
spec:
  accessModes:
  - ReadWriteMany
  storageClassName: file-store
  resources:
    requests:
      storage: 10Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: kickstar-backend
spec:
  replicas: 1
  revisionHistoryLimit: 4
  selector:
    matchLabels:
      app: backend
  template:
    metadata:
      labels:
        app: backend
    spec:
      containers:
      - image: backend:latest
        name: backend
        ports:
        - containerPort: 1026
        volumeMounts:
        - mountPath: /app/wwwroot
          name: file-store
      volumes:
      - name: file-store
        persistentVolumeClaim:
          claimName: file-store-pvc

问题分析

从Pod事件中可以看到两个关键错误:

  1. An operation with the given volume key ... already exists:表示针对该Filestore卷的操作正在进行中,可能是之前的挂载请求未正常完成,导致资源被占用。
  2. context deadline exceeded:挂载操作超时,说明节点无法在规定时间内完成与Filestore实例的连接或挂载操作,大概率是网络连通性问题或Filestore实例异常。

同时注意到PVC请求的是10Gi存储,但实际创建的PV是1Ti,这是因为Filestore的最小存储容量为1Ti,属于正常现象,不会导致挂载失败。

解决步骤

1. 检查Filestore实例状态

确认自动创建的Filestore实例是否正常运行,且与GKE节点处于同一个VPC网络:

gcloud filestore instances list --region asia-southeast1

查看实例状态应为READY,网络配置与节点所在VPC一致。

2. 验证节点与Filestore的网络连通性

登录到Pod所在的GKE节点,测试Filestore实例的NFS端口(2049)是否可达:

# 获取Filestore实例的IP
FILESTORE_IP=$(gcloud filestore instances describe <instance-name> --region asia-southeast1 --format="value(networks.ipAddresses[0])")
# 测试连通性
nc -zv $FILESTORE_IP 2049

如果连接失败,检查VPC防火墙规则是否允许节点与Filestore之间的NFS流量(默认Filestore会创建允许所有内部IP访问2049端口的规则,确认未被修改)。

3. 检查Filestore CSI驱动状态

确认GKE集群中的Filestore CSI驱动Pod是否正常运行:

kubectl get pods -n kube-system -l app=filestore-csi-node
kubectl get pods -n kube-system -l app=filestore-csi-controller

所有Pod应处于Running状态,若有异常,重启对应的Pod或重新部署CSI驱动。

4. 清理残留挂载操作

如果是之前的挂载操作卡住导致资源占用,尝试删除PVC并重新创建:

# 删除PVC(会自动删除对应的PV和Filestore实例)
kubectl delete pvc file-store-pvc
# 重新应用PVC配置
kubectl apply -f <your-pvc-config-file>.yaml

之后重启Deployment,让Pod重新绑定新的PV。

5. 重启节点kubelet服务

如果节点上的kubelet进程异常,导致挂载操作无法执行,登录节点重启kubelet:

sudo systemctl restart kubelet

6. 确认StorageClass网络参数

检查StorageClass中的network参数是否为节点所在的网络名称,确保Filestore实例创建在正确的网络中。

内容的提问来源于stack exchange,提问作者Vuong Tran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 17:19:53