GKE上带持久化的Ignite集群重部署出现FailedMount等错误求助
解决GKE上Ignite集群持久化部署后的两类错误
问题场景
在GKE部署2节点Ignite集群,已启用默认区域持久化并配置PVC,测试数据持久化时出现以下问题:
- 首次部署且未写入缓存数据时,重新部署集群正常
- 写入一条缓存数据后,销毁集群重新部署(不删除PVC)时触发两类错误
错误1:容器创建阶段FailedMount
MountVolume.SetUp failed for volume "pvc-1a83b66a-8be9-49bc-b015-4e4b5b237e37" : applyFSGroup failed for vol projects/UNSPECIFIED/zones/xxx/disks/pvc-1a83b66a-8be9-49bc-b015-4e4b5b237e37: readdirent /var/lib/kubelet/pods/ae17e40c-c924-4d38-a232-c4b853ffb1ae/volumes/kubernetes.io~csi/pvc-1a83b66a-8be9-49bc-b015-4e4b5b237e37/mount: input/output error
错误2:节点2 CrashLoopBackOff
节点2日志输出:
Failed to start manager: GridManagerAdapter [enabled=true, name=o.a.i.i.managers.discovery.GridDiscoveryManager] class org.apache.ignite.IgniteCheckedException: Failed to start SPI: java.net.ConnectException: Connection refused (Connection refused)]
关联StatefulSet PVC配置片段
containers: - name: ignite-node image: "gridgain/community:8.8.19" imagePullPolicy: Always resources: {{- toYaml .Values.resources | nindent 12 }} env: - name: OPTION_LIBS value: ignite-kubernetes,ignite-rest-http,control-center-agent - name: CONFIG_URI value: file:///opt/ignite/config/node-configuration.xml - name: JVM_OPTS value: "-DIGNITE_WAL_MMAP=false \ -DIGNITE_WAIT_FOR_BACKUPS_ON_SHUTDOWN=true \ - ...." ports: - containerPort: ... volumeMounts: - mountPath: opt/ignite/config name: config-vol - mountPath: opt/ignite/work name: work-vol - mountPath: opt/ignite/wal name: wal-vol - mountPath: opt/ignite/walarchive name: walarchive-vol volumeClaimTemplates: - metadata: name: work-vol spec: accessModes: [ "ReadWriteOnce" ] resources: requests: storage: "1G" - metadata: name: wal-vol spec: accessModes: [ "ReadWriteOnce" ] resources: requests: storage: "1G" - metadata: name: walarchive-vol spec: accessModes: [ "ReadWriteOnce" ] resources: requests: storage: "1G"
修复方案
针对错误1(FailedMount)
- 检查并修复文件系统
I/O错误多由磁盘文件系统损坏导致,可通过临时Pod挂载磁盘执行检查:kubectl run fsck-pod --image=ubuntu --rm -it --attach --volume pvc=pvc-1a83b66a-8be9-49bc-b015-4e4b5b237e37 --volume mountPath=/mnt # 在临时Pod内执行(替换为实际磁盘设备名) fsck /dev/sdb - 显式设置FSGroup
在StatefulSet的securityContext中指定Ignite运行用户ID(默认1000),避免Kubernetes自动修改权限触发错误:spec: template: spec: securityContext: fsGroup: 1000 - 切换RWX存储类
GCE Persistent Disk默认是ReadWriteOnce,多节点场景易出现挂载冲突。改用Filestore等支持ReadWriteMany的存储类,确保PVC可被多节点正常挂载。
针对错误2(节点2连接拒绝)
- 清理持久化目录残留数据
Ignite重启时会读取work目录的旧集群状态,残留信息会导致发现失败。重新部署前清理对应PVC的work目录:kubectl run cleanup-pod --image=ubuntu --rm -it --attach --volume pvc=<work-vol-pvc-name> --volume mountPath=/opt/ignite/work # 在Pod内执行 rm -rf /opt/ignite/work/* - 验证Kubernetes发现SPI配置
确认node-configuration.xml中发现SPI配置正确,节点可通过服务名找到集群:<bean class="org.apache.ignite.spi.discovery.tcp.TcpDiscoverySpi"> <property name="ipFinder"> <bean class="org.apache.ignite.spi.discovery.tcp.ipfinder.kubernetes.TcpDiscoveryKubernetesIpFinder"> <property name="namespace" value="your-namespace"/> <property name="serviceName" value="ignite-service"/> </bean> </property> </bean> - 添加节点启动延迟
在StatefulSet中为节点添加启动延迟,确保主节点完全就绪后再启动从节点:containers: - name: ignite-node lifecycle: postStart: exec: command: ["sleep", "30"]
内容的提问来源于stack exchange,提问作者Vincent Y
相关产品推荐
相关产品推荐

