You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

三节点Consul集群部署报错No cluster Leader 仅单节点UI可访问

三节点Consul集群"No cluster Leader"故障排查

问题现象

  • 搭建三节点Consul服务器集群过程中,仅1台服务器可正常通过Consul UI访问,其余主机访问Consul UI时返回No cluster Leader错误
    Consul UI报错截图

异常节点错误日志

故障Consul Server节点输出日志如下:

2022-07-06T14:54:29.925Z [WARN]  agent: [core]grpc: addrConn.createTransport failed to connect to {dc1-17.99.211.49:8300 lapp116.dc1 <nil> 0 <nil>}. Err: connection error: desc = "transport: Error while dialing dial tcp <IP-1>:0-><Ip-2>:8300: operation was canceled". Reconnecting...
2022-07-06T14:54:41.537Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:55:05.054Z [ERROR] agent: Coordinate update error: error="No cluster leader"
2022-07-06T14:55:06.052Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:55:28.927Z [ERROR] agent: Coordinate update error: error="No cluster leader"
2022-07-06T14:55:31.513Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:56:02.307Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:56:03.070Z [ERROR] agent: Coordinate update error: error="No cluster leader"
2022-07-06T14:56:26.165Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:56:35.031Z [ERROR] agent: Coordinate update error: error="No cluster leader"
2022-07-06T14:56:55.459Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:57:07.616Z [ERROR] agent: Coordinate update error: error="No cluster leader"
2022-07-06T14:57:27.686Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:57:35.128Z [ERROR] agent: Coordinate update error: error="No cluster leader"
2022-07-06T14:57:58.915Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:58:03.708Z [ERROR] agent: Coordinate update error: error="No cluster leader"
2022-07-06T14:58:22.158Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader"
2022-07-06T14:58:29.168Z [ERROR] agent: Coordinate update error: error="No cluster leader"
2022-07-06T14:58:50.469Z [ERROR] agent.anti_entropy: failed to sync remote state: error="No cluster leader

根因分析

从日志第一条gRPC连接8300端口失败的告警可以判断,故障节点无法和集群其他Server节点建立Raft协议通信,Raft选主流程无法推进,最终持续抛出无集群Leader的错误。常见触发原因如下:

  • 网络层拦截:节点间安全组、系统防火墙未放开Consul Server通信所需端口,8300(Raft通信)、8301(LAN Gossip)、8302(WAN Gossip)端口不通,导致节点间无法交换Raft投票信息,达不到3节点集群选主需要的2票法定人数
  • 启动配置错误:节点配置中bootstrap_expect参数值不为3、retry_join地址列表未覆盖全部节点、节点未标记为Server角色,都会导致集群无法完成初始化选主
  • 本地状态损坏:故障节点之前启动残留的Raft持久化文件损坏,无法正常参与Raft投票流程

修复步骤

  1. 排查端口连通性
    在每个故障节点上执行端口探测命令,验证和其余两个节点的通信是否正常:
    # 替换<peer-ip>为其他节点的实际IP
    nc -zv <peer-ip> 8300
    nc -zv <peer-ip> 8301
    nc -zv <peer-ip> 8302
    
    若探测失败,逐一调整云平台安全组、主机iptables/firewalld规则,放开三个Server节点间上述端口的双向访问权限。
  2. 核对节点配置
    逐一检查三个Server节点的Consul配置文件,确认以下参数完全符合要求:
    • server = true:三个节点均为Server角色
    • bootstrap_expect = 3:所有节点该参数值统一为3
    • datacenter:所有节点数据中心名称一致
    • retry_join:地址列表包含全部三个节点的业务通信IP,不得使用127.0.0.1等回环地址
  3. 重置故障节点状态
    网络和配置校验完成后,先停止所有节点的Consul服务,仅在故障节点上删除数据目录下的raft文件夹(数据目录路径参考配置中data_dir参数,默认路径通常为/opt/consul/data),禁止删除当前可正常访问的节点上的Raft数据,避免原有集群状态丢失。
    清理完成后先启动原可正常访问的节点,等待10秒后依次启动两个故障节点。
  4. 集群状态校验
    集群启动完成后,在任意节点执行以下命令验证状态:
    # 查看节点列表,确认三个节点均为alive状态
    consul members
    # 查看Raft集群状态,确认存在1个Leader、2个Follower
    consul operator raft list-peers
    
    状态正常后,访问所有节点的Consul UI将不再出现无Leader报错。

内容的提问来源于stack exchange,提问作者Durga Deep

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 15:33:20