Elasticsearch跨集群快照恢复失败:索引无法分配且文件缺失求助
问题
在2节点集群1中通过curl创建快照后,尝试在单节点集群2中通过curl恢复快照时,集群2无法分配全部索引。快照文件夹中确实存在文件snap-hb19TYAWQ1SFZvTzd-2AkA.dat,通过_cluster/allocation/explain排查到以下错误:
nested: IndexShardRestoreFailedException[failed to restore snapshot [05_01_24/hb19TYAWQ1SFZvTzd-2AkA]];
nested: SnapshotMissingException[[my_repo:05_01_24/hb19TYAWQ1SFZvTzd-2AkA] is missing];
nested: NoSuchFileException[blob object [snap-hb19TYAWQ1SFZvTzd-2AkA.dat] not found];
- manually close or delete the index [tasklist-task-8.2.3_2023-11-30] in order to retry to restore the snapshot again or use the reroute API to force the allocation of an empty primary shard
解决方法
方案1:关闭问题索引后重新恢复快照
- 关闭出错的索引:
curl -X POST "http://<集群2节点IP>:<端口>/_close/tasklist-task-8.2.3_2023-11-30"
- 重新执行快照恢复命令(替换为实际的恢复参数):
curl -X POST "http://<集群2节点IP>:<端口>/_snapshot/my_repo/05_01_24/_restore" -H "Content-Type: application/json" -d '{ "indices": "tasklist-task-8.2.3_2023-11-30" }'
方案2:删除问题索引后重新恢复快照
若关闭索引后仍无法恢复,直接删除索引再重试:
- 删除索引:
curl -X DELETE "http://<集群2节点IP>:<端口>/tasklist-task-8.2.3_2023-11-30"
- 重新执行上述快照恢复命令
方案3:使用reroute API强制分配空主分片
此操作会丢失对应分片的数据,仅在快照无法恢复时使用:
- 先通过以下命令查看目标索引的分片编号:
curl "http://<集群2节点IP>:<端口>/_cat/shards/tasklist-task-8.2.3_2023-11-30?v"
- 执行强制分配命令(替换
<分片编号>和<集群2节点名称>):
curl -X POST "http://<集群2节点IP>:<端口>/_cluster/reroute" -H "Content-Type: application/json" -d '{ "commands": [ { "allocate_empty_primary": { "index": "tasklist-task-8.2.3_2023-11-30", "shard": <分片编号>, "node": "<集群2节点名称>", "accept_data_loss": true } } ] }'
额外排查要点
- 确认集群2的快照仓库
my_repo配置路径正确,且Elasticsearch进程对该路径有读写权限。 - 检查快照文件夹的完整性:除
snap-hb19TYAWQ1SFZvTzd-2AkA.dat外,需确保存在meta-hb19TYAWQ1SFZvTzd-2AkA.dat等元数据文件,缺失元数据会导致快照无法被识别。 - 验证集群2的Elasticsearch版本与集群1创建快照的版本兼容:主版本号必须一致,小版本尽量保持接近,跨大版本恢复可能出现兼容性问题。
内容的提问来源于stack exchange,提问作者Epic555
相关产品推荐
相关产品推荐

