关于MariaDB Galera集群节点离线复归的同步速度、性能影响与请求可靠性的技术问询
Hey Bert, let's break down your questions about MariaDB Galera Cluster step by step—this is a super common concern for anyone running clustered databases, so you're not alone here. I'll cover each part clearly based on real-world Galera deployments:
1. How Fast Does Data Sync When a Node Rejoins?
The speed depends entirely on how much data the node missed while offline, and which sync method Galera uses:
- Incremental State Transfer (IST): If the offline period is short enough that the rejoining node can still access the incremental write logs (stored in the cluster's global cache,
gcache), it'll use IST. This is fast—think seconds to minutes, depending on the volume of writes that happened while it was offline. Since your nodes are on the same low-latency VLAN, network won't be a bottleneck here. - State Snapshot Transfer (SST): If the node was offline too long and the incremental logs are gone (gcache filled up or rotated), Galera will trigger a full SST—basically a full copy of the entire dataset from a "donor" node in the cluster. Speed here depends on your dataset size: a 50GB dataset on SSDs might take 10-30 minutes, while larger datasets on mechanical disks could take hours. You can speed this up by using faster SST methods like
rsyncinstead ofmysqldump(configure viawsrep_sst_method).
2. Will Syncing Impact Other Nodes' Performance?
Again, this splits into IST and SST scenarios:
- IST: Minimal impact. The donor node just sends a small batch of incremental logs, so it'll barely notice the extra load—your app can keep running normally on all nodes.
- SST: The donor node will have extra disk I/O (reading the full dataset) and network usage. On a low-latency VLAN, network isn't an issue, but if the donor is a busy production node, you might see a small dip in query performance (especially with disk-bound workloads). To mitigate this, you can:
- Configure a dedicated donor node (via
wsrep_sst_donor) that's not handling critical app traffic. - Schedule SST during off-peak hours if possible (though Galera triggers it automatically when a node rejoins).
- Use SSDs on all nodes to speed up disk I/O during SST.
- Configure a dedicated donor node (via
Important note: Even during SST, the donor node doesn't stop processing requests—it just runs the sync in the background alongside normal operations.
3. What Happens if the App Queries the Rejoining Node?
This is the critical part you're worried about, and the answer depends on how your app connects to the cluster:
- If you're using a load balancer (recommended): Most load balancers (like HAProxy) can be configured to check Galera's node status via
SHOW STATUS LIKE 'wsrep_ready'. If the rejoining node is still syncing,wsrep_readywill beOFF, and the load balancer will stop sending requests to it entirely. Your app will only hit fully synced nodes, so no errors here—100% reliable. - If your app connects directly to individual nodes: If the app sends a request to the rejoining node while it's in the "Joiner" state (not fully synced), you will get errors. The node will either:
- Reject write operations with an error like
ERROR 1047 (08S01): Unknown command(sincewsrep_readyis OFF and it can't participate in cluster writes). - Return stale, outdated data for read operations (if you haven't enabled
wsrep_sync_waitto force reads to wait for sync).
There's no automatic "fallback" to other nodes here—your app would need retry logic or a client driver that supports cluster-aware load balancing (like MariaDB's Connector/J with load balancing enabled) to redirect requests to healthy nodes.
- Reject write operations with an error like
Quick Recommendations to Avoid Issues
- Always use a load balancer with Galera health checks (monitor
wsrep_readyandwsrep_cluster_size). - Increase
gcache.size(viawsrep_provider_options) to store more incremental logs—this reduces the chance of triggering a full SST when nodes rejoin after short offline periods. - Use cluster-aware client drivers or implement retry logic in your app to handle temporary node unavailability.
备注:内容来源于stack exchange,提问作者Bert

