Redis集群同步写入模式下为何仍可能出现数据丢失?
Great question! Let's break down why even when using the WAIT command for synchronous writes, Redis Cluster can still lose data in complex failure scenarios. First, let's recap the synchronous write flow you outlined to set the stage:
- Client Z sends data D to master node B
- Master B replicates D to its replicas B1, B2, B3
- Master B waits for ack responses from specified replicas (via
WAIT) - Master B returns success to client Z
While this flow cuts down on write loss risk, it doesn't guarantee strong consistency. Here are the key reasons this can happen:
1. WAIT only confirms in-memory receipt, not persistence
When a replica sends an ack to the master, it only means the replica has loaded the write into its memory—it doesn’t ensure the data was saved to disk (via RDB or AOF).
Imagine this scenario: Master B gets acks from B1 and B2, then crashes suddenly. B1 restarts shortly after, but since it hadn’t persisted that write to disk before shutting down, it loses the data on boot. If B1 is elected as the new master, that write is gone for good.
2. Network partitions ("split brain")
Let’s say your cluster splits into two isolated network partitions:
- Partition 1: Master B, client Z, replicas B1 and B2 (all can communicate with each other)
- Partition 2: Replica B3 (cut off from partition 1)
Master B processes the write, gets acks from B1/B2, and is about to send a success response to Z—then the connection between partition 1 and the rest of the cluster drops.
In partition 2, nodes detect that B is unreachable and elect B3 as the new master. B3 never received the original write. Client Z, not getting a response from B, retries the write to B3. When the network heals, original master B becomes a replica of B3 and syncs B3’s data—overwriting the original write and making it lost.
3. Election of replicas with incomplete data
Suppose you set WAIT 1, so master B only waits for one replica to ack. B sends the write to B1, B2, B3, gets an ack from B1, then both B and B1 crash immediately.
The cluster now has to elect a new master from B2 and B3—neither of which received that write. The elected master will have no record of the data, and the write is lost.
Bottom Line
Redis Cluster is built to prioritize availability and performance over strong consistency (following the CAP theorem's AP tradeoff). The WAIT command reduces write loss risk by ensuring replicas have the data in memory, but it can’t cover every edge case where replicas with incomplete data get elected as masters, or persistence gaps lead to data loss.
内容的提问来源于stack exchange,提问作者Vibhor Verma

