如何在Kubernetes中备份分片MongoDB?GCP+Bitnami Operator恢复求助
针对Bitnami MongoDB分片集群(GCP环境)的备份恢复可行方案
一、使用Operator原生Backup/Restore CRD(优先推荐)
Bitnami MongoDB Operator自带适配分片集群的MongoDBBackup和MongoDBRestore自定义资源,是最贴合当前环境的方案:
备份步骤
- 创建
MongoDBBackup资源,指定GCS作为存储:apiVersion: mongodb.bitnami.com/v1alpha1 kind: MongoDBBackup metadata: name: sharded-mongo-backup namespace: your-namespace spec: mongodbResourceRef: name: your-sharded-mongo-cluster storage: type: gcs gcs: bucket: your-gcs-bucket-name secretRef: name: gcs-service-account-secret # 提前创建含GCP服务账号密钥的Secret - 查看备份状态:
kubectl get mongodbbackups -n your-namespace,status.phase显示Completed即为备份完成。
恢复步骤
- 确保目标集群(新建或现有)的分片数、副本集成员数与备份源完全一致,且集群正常运行。
- 创建
MongoDBRestore资源指向GCS备份文件:apiVersion: mongodb.bitnami.com/v1alpha1 kind: MongoDBRestore metadata: name: sharded-mongo-restore namespace: your-namespace spec: mongodbResourceRef: name: your-target-sharded-cluster backupSource: gcs: bucket: your-gcs-bucket-name path: backups/sharded-mongo-cluster/<backup-timestamp-dir> # 替换为实际备份目录 secretRef: name: gcs-service-account-secret - 监控恢复状态:
kubectl get mongodbrestores -n your-namespace,完成后验证数据:kubectl exec -it <mongos-pod-name> -n your-namespace -- mongosh # 执行show dbs、db.collection.find()等命令验证数据完整性
二、手动备份恢复(Operator CRD方式失效时备选)
备份操作
- 临时锁定集群保证一致性:
kubectl exec -it <mongos-pod-name> -- mongosh --eval "sh.stopBalancer(); db.fsyncLock();" - 逐个备份分片主节点与配置服务器:
# 备份分片1主节点 kubectl exec -it <shard1-primary-pod> -- mongodump --uri="mongodb://localhost:27017" --db=your-db --out=/tmp/shard1-backup kubectl cp <shard1-primary-pod>:/tmp/shard1-backup ./local-shard1-backup # 备份配置服务器主节点 kubectl exec -it <configsvr-primary-pod> -- mongodump --uri="mongodb://localhost:27017" --config --out=/tmp/config-backup kubectl cp <configsvr-primary-pod>:/tmp/config-backup ./local-config-backup - 解锁集群:
kubectl exec -it <mongos-pod-name> -- mongosh --eval "db.fsyncUnlock(); sh.startBalancer();" - 上传备份到GCS:
gsutil cp -r ./local-shard1-backup gs://your-gcs-bucket/sharded-backup/<timestamp>/shard1 gsutil cp -r ./local-config-backup gs://your-gcs-bucket/sharded-backup/<timestamp>/configsvr
恢复操作
- 停止目标集群Balancer:
kubectl exec -it <target-mongos-pod> -- mongosh --eval "sh.stopBalancer();" - 逐个恢复分片与配置服务器数据:
# 恢复分片1 kubectl cp ./local-shard1-backup <target-shard1-primary-pod>:/tmp/shard1-backup kubectl exec -it <target-shard1-primary-pod> -- mongorestore --uri="mongodb://localhost:27017" /tmp/shard1-backup # 恢复配置服务器 kubectl cp ./local-config-backup <target-configsvr-primary-pod>:/tmp/config-backup kubectl exec -it <target-configsvr-primary-pod> -- mongorestore --uri="mongodb://localhost:27017" --config /tmp/config-backup - 更新分片路由(若目标集群地址变化):
kubectl exec -it <target-mongos-pod> -- mongosh --eval " sh.addShard('shard1/<target-shard1-pod-0>:27017,<target-shard1-pod-1>:27017'); # 按需添加其他分片 " - 启动Balancer并验证:
kubectl exec -it <target-mongos-pod> -- mongosh --eval "sh.startBalancer();" # 执行数据验证命令确认恢复结果
三、常见恢复失败排查点
- GCP权限检查:确保服务账号Secret拥有GCS读写权限,可通过
gsutil ls gs://your-bucket测试。 - 集群规格匹配:目标集群的分片数、副本集成员数必须与备份源完全一致。
- 备份文件完整性:检查GCS中的备份目录是否包含完整的
metadata.json及数据文件,对比本地备份结构。 - 集群状态:恢复前确认目标集群所有Pod处于
Running状态,无CrashLoopBackOff异常。
内容的提问来源于stack exchange,提问作者jayanthi vikas
相关产品推荐
相关产品推荐

