Mariadb Galera 10.5.13-16节点崩溃求助:附日志及环境信息
Galera节点崩溃排查(信号6+Boost异常)
环境概述
- 集群架构:2个Galera节点 + 1个仲裁节点
- 流量代理:2台HAProxy
- 崩溃节点:节点1
- 数据库版本:10.5.13-MariaDB-1:10.5.13+maria~focal
关键日志信息提取
# 握手失败警告 2023-01-03 12:08:55 0 [Warning] WSREP: Handshake failed: peer did not return a certificate 2023-01-03 12:08:56 0 [Warning] WSREP: Handshake failed: http request # 致命异常 terminate called after throwing an instance of 'boost::wrapexcept<std::system_error>' what(): remote_endpoint: Transport endpoint is not connected # 崩溃信号触发 230103 12:08:56 [ERROR] mysqld got signal 6 ;
栈回溯显示崩溃链终止于libpthread.so.0与libgalera_smm.so,说明问题发生在Galera集群通信线程中。
排查与修复方向
1. 集群SSL通信配置校验
- 检查所有节点
wsrep_provider_options中的SSL参数:ssl_cert、ssl_key、ssl_ca路径是否正确,文件权限是否允许mysql用户读取 - 确认所有节点使用同一CA签发的有效证书,无过期或不匹配情况
- 排查HAProxy配置:是否错误将HTTP流量转发到Galera集群通信端口(默认4567),而非MySQL服务端口(默认3306)
2. 线程资源与pthread库校验
- 日志显示
thread_count=106超过max_threads=102,线程资源耗尽触发异常:- 降低
max_connections参数,避免线程数突破系统限制 - 检查系统线程上限:执行
ulimit -u查看用户线程配额,不足则修改/etc/security/limits.conf调整
- 降低
- 验证pthread库完整性:
若库文件损坏,重新安装对应系统包ldd /usr/sbin/mariadbd | grep libpthread dpkg -S /lib/x86_64-linux-gnu/libpthread.so.0
3. 集群网络稳定性排查
remote_endpoint: Transport endpoint is not connected表明节点间连接异常中断:- 检查节点间网络带宽、延迟与丢包率,确认无网络波动
- 验证防火墙/安全组是否放行Galera所需端口(4567-4569)
- 确认仲裁节点网络连通性,未出现集群隔离情况
4. 核心文件深度分析
利用生成的core文件定位具体崩溃点:
gdb /usr/sbin/mariadbd /var/lib/mysql/core bt full
通过栈回溯确认是否为Galera或MariaDB已知缺陷,针对性修复
临时恢复步骤
- 停止节点1的MariaDB服务
- 执行
mariadbd --wsrep-recover恢复集群状态 - 若恢复失败,从正常节点拷贝完整数据目录(确保节点处于一致状态)后重新加入集群
内容的提问来源于stack exchange,提问作者Theo Cerutti
相关产品推荐
相关产品推荐

