CockroachDB跨双数据中心灾难恢复方案咨询:单数据中心完全下线时的集群可用性保障
Let's break down practical solutions for both your deployment scenarios, focusing on keeping the cluster operational when an entire datacenter goes offline, plus seamless automatic recovery once connectivity is restored.
Scenario 1: 3 Nodes Across 2 Datacenters
The core requirement here is ensuring surviving nodes can form a Raft quorum (majority of replicas per Raft group) when one datacenter goes down. For 3 total replicas, the quorum is 2 nodes—so we need to structure the cluster to guarantee this majority exists in one datacenter.
Step-by-Step Configuration
Tag Nodes with Locality:
When starting each node, add a locality flag to tie it to its datacenter:# For nodes in DC1 cockroach start --locality=datacenter=dc1 --join=<node1-addr>,<node2-addr>,<node3-addr> ... # For the single node in DC2 cockroach start --locality=datacenter=dc2 --join=<node1-addr>,<node2-addr>,<node3-addr> ...We’ll place 2 nodes in DC1 and 1 in DC2—this ensures DC1 has enough nodes to form a quorum if DC2 goes offline.
Adjust Zone Configuration:
Modify the default zone to enforce replica distribution and set a survival goal for datacenter outages:ALTER RANGE default CONFIGURE ZONE USING num_replicas = 3, constraints = '{+datacenter=dc1: 2, +datacenter=dc2: 1}', -- Lock 2 replicas to DC1, 1 to DC2 survival_goal = 'zone'; -- Ensures cluster survives a single zone/datacenter outage
Behavior During & After Outage
- When DC2 goes offline: DC1’s 2 nodes form a valid quorum, so the cluster remains fully operational (supports both reads and writes).
- When connectivity is restored: The DC2 node will automatically rejoin the cluster, sync up with the latest data, and the cluster will restore its original 2/1 replica balance without manual intervention.
Note: If you prioritize DC2 for outage survival instead, swap the constraint values to place 2 replicas in DC2 and 1 in DC1.
Scenario 2: 6 Nodes (3 per Datacenter)
Your goal is to survive a full outage of either datacenter, with automatic recovery post-reconnection. The main challenge here is that a 3-node surviving cluster needs to meet Raft quorum requirements—let’s address this with targeted configurations.
Key Limitation to Fix
With the default 3-replica setup, spreading replicas evenly (2 per DC) means a full DC outage leaves only 1 replica per group, which can’t form a quorum. We need to adjust replica counts and distribution to fix this.
Recommended Solution 1: Asymmetric Replica Distribution (Survivable for One DC)
If you can prioritize one datacenter (e.g., DC1) for full outage survival, configure the cluster like this:
Tag Nodes with Locality: Same as Scenario 1, tag each node with
datacenter=dc1ordatacenter=dc2.Adjust Zone Configuration:
ALTER RANGE default CONFIGURE ZONE USING num_replicas = 5, -- Quorum becomes 3 nodes (5//2 + 1) constraints = '{+datacenter=dc1: 3, +datacenter=dc2: 2}', -- 3 replicas in DC1, 2 in DC2 survival_goal = 'datacenter';
Behavior
- When DC2 goes offline: DC1’s 3 nodes meet the quorum requirement, so the cluster stays fully operational.
- When connectivity is restored: DC2’s nodes rejoin, sync missing data, and the cluster restores the 3/2 replica balance automatically.
Solution 2: Survivability for Either DC
If you need both datacenters to be able to survive a full outage, you have two options:
- Add Nodes: Increase each datacenter to 4 nodes (total 8 nodes), then set
num_replicas=5with 3 replicas in each DC. This way, if either DC goes down, the remaining 3 nodes meet the quorum of 3. - Multi-Cluster with CDC: Deploy two independent clusters (one per DC) and enable Change Data Capture (CDC) for bidirectional sync. If one DC goes offline, the other cluster continues operating. When connectivity is restored, CDC syncs any missing changes automatically—this is the only way to achieve active-active survivability for two symmetric datacenters.
Automatic Recovery Notes
For all configurations, once the offline datacenter reconnects:
- Nodes will automatically rejoin the cluster.
- Raft will sync missing log entries to restore replica consistency.
- The cluster will rebalance replicas back to the configured distribution without manual work.
内容的提问来源于stack exchange,提问作者Madhu

