Elasticsearch节点重启场景下同步写入机制的技术问询
Great question! Let's walk through exactly how Elasticsearch handles this scenario, based on its core replication mechanics:
先回顾默认的同步写入逻辑
First, to set the stage: when you're writing to an index where replicas are spread across nodes 1 and 2, the default synchronous write flow (where the client gets a success response only after both primary and all replica shards confirm the write) works like this:
- The write request hits the coordinating node, which routes it to the primary shard's node.
- The primary shard writes the data to its local store and translog.
- It then replicates the write to all assigned replica shards.
- Only when all replicas confirm they've successfully written the data does the primary send a success response back to the client.
当其中一个副本节点异常时的变化
Now, if one of the nodes (say node 1) goes offline or becomes unresponsive:
Immediate cluster detection
- Elasticsearch's cluster state monitoring will quickly mark the replica shard on node 1 as
UNASSIGNEDorOFFLINE. The cluster will start trying to either:- Wait for node 1 to come back online, or
- Initiate a reallocation of that replica shard to another available node (if resources allow).
- Elasticsearch's cluster state monitoring will quickly mark the replica shard on node 1 as
Impact on ongoing write requests
- If you're using the strict synchronous mode (waiting for all shards to confirm), your write requests will block temporarily. The primary shard will still write the data to its own store and translog, but it can't get confirmation from the missing replica.
- The primary will keep retrying the replication to the missing replica (or the newly allocated one once it's spun up) until it gets that confirmation. Only then will it send a success response to your client.
Data consistency guarantees
- Don't worry about data loss here: all unconfirmed writes are stored in the primary's translog. Once the missing replica is back online (or a new replica is created and caught up), the primary will replay all pending operations from the translog to sync the replica up to date.
Workarounds if you need immediate writes
- If you can tolerate temporarily writing only to the primary (with the replica catching up later), you can adjust the
wait_for_active_shardsparameter in your write request. For example, setting it to1means the primary only needs to confirm the write, ormajorityif you want a quorum of shards to confirm. This will let writes succeed immediately even if some replicas are offline.
- If you can tolerate temporarily writing only to the primary (with the replica catching up later), you can adjust the
Note: Changing
wait_for_active_shardstrades off consistency for availability, so make sure this aligns with your application's requirements.
内容的提问来源于stack exchange,提问作者Shivam Dixit

