MariaDB 10.4 Galera-4集群节点重同步后性能骤降求助
MariaDB 10.4+Galera-4集群重同步后性能骤降问题分析与解决
问题描述
基于MariaDB 10.4和Galera-4搭建的主主集群原本运行正常,在部分节点间出现网络问题后,节点重新连接并完成状态重同步,但整个集群性能急剧下降,目前仅能通过从捐赠节点重建整个集群恢复性能。
相关日志
2023-03-19 1:23:20 0 [Note] WSREP: ####### Adjusting cert position: 61149358 -> 61149359 2023-03-19 1:23:20 0 [Note] WSREP: Service thread queue flushed. 2023-03-19 1:23:20 0 [Note] WSREP: Lowest cert index boundary for CC from ist: 61149189 2023-03-19 1:23:20 0 [Note] WSREP: Min available from gcache for CC from ist: 60988106 2023-03-19 1:23:20 0 [Note] WSREP: Receiving IST...100.0% (545/545 events) complete. 2023-03-19 1:23:21 2 [Note] WSREP: ================================================ View: id: 6caf7137-be5b-11ed-952d-8e2998c314a7:61149359 status: primary protocol_version: 4 capabilities: MULTI-MASTER, CERTIFICATION, PARALLEL_APPLYING, REPLAY, ISOLATION, PAUSE, CAUSAL_READ, INCREMENTAL_WS, UNORDERED, PREORDERED, STREAMING, NBO final: no own_index: 2 members(4): 0: 38ac2156-c3e1-11ed-b954-ba79e6d7687d, DB1 1: 637c05dc-c17f-11ed-b3f2-0e364b289bef, DB3 2: 667386bc-c5db-11ed-ad68-ee03169ac5b9, DB4 3: c9c34048-c3e0-11ed-8cbc-13957f2241e4, DB2 ================================================ 2023-03-19 1:23:21 2 [Note] WSREP: Server status change initialized -> joined 2023-03-19 1:23:21 2 [Note] WSREP: wsrep_notify_cmd is not defined, skipping notification. 2023-03-19 1:23:21 2 [Note] WSREP: wsrep_notify_cmd is not defined, skipping notification. 2023-03-19 1:23:21 2 [Note] WSREP: Draining apply monitors after IST up to 61149359 2023-03-19 1:23:21 2 [Note] WSREP: IST received: 6caf7137-be5b-11ed-952d-8e2998c314a7:61149359 2023-03-19 1:23:21 2 [Note] WSREP: Lowest cert index boundary for CC from sst: 61149189 2023-03-19 1:23:21 2 [Note] WSREP: Min available from gcache for CC from sst: 60988107 2023-03-19 1:23:21 0 [Note] WSREP: 2.0 (DB4): State transfer from 3.0 (DB2) complete. 2023-03-19 1:23:21 0 [Note] WSREP: Shifting JOINER -> JOINED (TO: 61149618) 2023-03-19 1:23:21 0 [Note] WSREP: Processing event queue:... 0.0% ( 0/260 events) complete. 2023-03-19 1:23:32 2 [Note] WSREP: Processing event queue:... 56.8% (208/366 events) complete. 2023-03-19 1:23:40 0 [Note] WSREP: Member 2.0 (DB4) synced with group. 2023-03-19 1:23:40 0 [Note] WSREP: Processing event queue:...100.0% (402/402 events) complete. 2023-03-19 1:23:40 0 [Note] WSREP: Shifting JOINED -> SYNCED (TO: 61149756) 2023-03-19 1:23:43 2 [Note] WSREP: Server DB4 synced with group. 2023-03-19 1:23:43 2 [Note] WSREP: Server status change joined -> synced 2023-03-19 1:23:43 2 [Note] WSREP: Synchronized with group, ready for connections 2023-03-19 1:23:43 2 [Note] WSREP: wsrep_notify_cmd is not defined, skipping notification. 2023-03-19 1:23:59 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61150337, -110 (Connection timed out) 2023-03-19 1:24:15 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61151132, -110 (Connection timed out) 2023-03-19 1:24:17 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61151205, -110 (Connection timed out) 2023-03-19 1:24:23 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61151478, -110 (Connection timed out) 2023-03-19 1:24:33 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61151724, -110 (Connection timed out) 2023-03-19 1:24:38 0 [Note] InnoDB: Buffer pool(s) load completed at 230319 1:24:38 2023-03-19 1:24:41 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61152114, -110 (Connection timed out) 2023-03-19 1:24:43 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61152261, -110 (Connection timed out) 2023-03-19 1:24:46 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61152364, -110 (Connection timed out) 2023-03-19 1:27:32 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61158126, -110 (Connection timed out) 2023-03-19 1:27:36 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61158269, -110 (Connection timed out) 2023-03-19 1:27:40 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61158385, -110 (Connection timed out) 2023-03-19 1:27:47 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61158486, -110 (Connection timed out) 2023-03-19 1:28:12 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61159789, -110 (Connection timed out) 2023-03-19 1:28:15 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61159889, -110 (Connection timed out) 2023-03-19 1:28:18 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61160056, -110 (Connection timed out) 2023-03-19 1:28:21 0 [Warning] WSREP: Failed to report last committed 6caf7137-be5b-11ed-952d-8e2998c314a7:61160171, -110 (Connection timed out)
当前wsrep配置
[galera] wsrep_on=ON wsrep_provider=/usr/lib/libgalera_smm.so wsrep_cluster_address="gcomm://xx.xx.xx.xxx,xxx.xx.xx.xxx,xx.xxx.xx.xxx,xx.xxx.xxx.xx" wsrep_sst_method=mariabackup wsrep_sst_auth=xxxxxxxxx:xxxxxxxxx binlog_format=row default_storage_engine=InnoDB innodb_autoinc_lock_mode=2 wsrep_cluster_name="xxxxxxx" wsrep_node_address="xx.xxx.xxx.xx" wsrep_node_name="DBX"
性能问题成因
- 网络问题未彻底解决:日志中持续出现
Failed to report last committed... Connection timed out,说明节点间网络仍存在不稳定,Galera组通信(GCS)频繁超时,集群需要不断重试通信、维护数据一致性,消耗大量CPU和带宽资源,直接拖垮性能。 - IST重同步后的元数据不一致:日志显示重同步时调整了证书位置(
Adjusting cert position),可能导致节点间的事务证书索引、GCache缓存出现局部不一致,集群在事务认证阶段需要额外的校验和回溯操作,增加了负载。 - 事务队列与并行应用异常:重同步后处理事件队列的过程中,可能存在未完全清理的旧事务残留,或者并行应用机制(PARALLEL_APPLYING)因网络波动出现异常,导致事务应用效率大幅降低。
解决方法
1. 彻底修复节点间网络问题
- 用
ping、traceroute测试节点间网络延迟和连通性,用nc -zv <节点IP> 4567/4568/4569验证Galera通信端口是否畅通 - 排查防火墙、安全组是否存在临时拦截规则,检查网络设备是否有丢包、带宽瓶颈
- 若为跨机房集群,确认专线或公网链路是否存在抖动、丢包率过高的情况
2. 优化Galera网络与同步参数
在[galera]配置段添加或修改以下参数,优化组通信和同步效率:
wsrep_provider_options="gcache.size=4G; gcs.fc_limit=16; gcs.fc_factor=0.5; evs.send_window=4096; evs.user_send_window=4096; evs.inactive_check_period=PT1S; evs.inactive_timeout=PT10S; evs.connect_timeout=PT10S" wsrep_slave_threads=8 # 根据CPU核心数调整,建议为核心数的2-4倍 wsrep_causal_reads=ON
gcache.size:调整GCache大小,减少SST触发概率evs.*参数:优化组通信超时和窗口大小,减少因网络波动导致的重试wsrep_slave_threads:提升事务并行应用能力
3. 清理异常节点的GCache并重新同步
- 将异常节点从集群中临时移除:修改
wsrep_cluster_address=gcomm://后重启节点 - 删除GCache文件(默认路径为
/var/lib/mysql/galera.cache,具体以实际配置为准) - 恢复
wsrep_cluster_address为原集群地址,重启节点触发IST同步
4. 增量重建异常节点(替代全量集群重建)
选择状态稳定的捐赠节点,对异常节点执行SST同步:
- 确保捐赠节点的
wsrep_sst_method=mariabackup配置正确,且备份权限正常 - 在异常节点上重启服务,自动触发从捐赠节点的SST同步,无需重建整个集群
内容的提问来源于stack exchange,提问作者uzurik
相关产品推荐
相关产品推荐

