Redis哨兵在脑裂恢复后的行为及多子网部署配置问询
Great question—this is a critical scenario for multi-subnet Redis deployments where you’ve intentionally allowed split-brain to maintain availability. Let’s walk through exactly what happens when your two subnets reconnect after being partitioned.
First, a quick recap of your setup to ground this: you have 4 Redis nodes (2 per subnet), configured so that if the inter-subnet link drops, each subnet can elect its own master and run independently. When the link comes back up, here’s how Sentinel handles the split-brain resolution:
Step 1: Sentinel Cluster Reconnects and Exchanges State
As soon as network connectivity is restored, all Sentinel nodes across both subnets will rediscover each other (via their configured sentinel monitor settings). They’ll immediately start sharing everything they know about the cluster:
- Which nodes are online
- Which master each Sentinel recognizes
- Replication status of all slaves
- Epoch values (the logical "version number" of a master, incremented every failover)
Step 2: Sentinel Consensus Picks the "Legitimate" Master
Sentinels use two key metrics to resolve the dual-master conflict:
- Epoch Value: The master with the higher epoch is automatically considered the valid one. This makes sense because the epoch increments every time a failover happens—so whichever subnet triggered a failover later (or had the last successful failover) will have the higher epoch.
- Run ID: If by some rare chance both masters have the same epoch, Sentinels will compare their unique
runid(generated when a node starts). The master with the lexicographically larger run ID wins.
Once consensus is reached, the Sentinel cluster will agree on which master is the "real" one and which is the rogue master from the split-brain.
Step 3: Demote the Rogue Master to Slave
The Sentinel cluster will send a SLAVEOF command to the rogue master, forcing it to become a slave of the legitimate master. At this point:
- The rogue master stops accepting write requests immediately
- It begins syncing data from the legitimate master. Any writes that happened on the rogue master during the split-brain will be overwritten by the legitimate master’s dataset. This is a key trade-off of allowing split-brain for availability—you have to accept that some data from the partitioned subnet may be lost.
Step 4: Cluster Resumes Normal Operation
Once the former rogue master finishes syncing (or starts syncing, depending on your persistence settings), the cluster is back to a single-master state. Sentinels will resume their normal monitoring duties, watching all nodes and ready to handle future failures.
Key Notes to Keep in Mind
- Quorum Configuration: Make sure your Sentinel
quorumis set so that each subnet can independently reach quorum to trigger a failover during a split. For example, if you have 4 Sentinels (2 per subnet), setting quorum to 2 lets each subnet elect a master when partitioned, and when reconnected, all 4 can quickly reach consensus. - Data Loss Risk: If both masters accepted writes during the split, the legitimate master’s data will overwrite the rogue’s. If this is a problem, consider adding application-level checks for write conflicts, or use Redis features like
min-slaves-to-writeto limit writes on a master that can’t reach enough slaves (though this would reduce availability during split-brain, which you’re trying to avoid). - Timing Parameters: Adjust
down-after-millisecondsandfailover-timeoutto fine-tune how quickly Sentinels detect failures and resolve split-brain. Shorter timeouts mean faster recovery, but increase the risk of false positives.
内容的提问来源于stack exchange,提问作者Nikolay Kuznetsov

