You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kraft模式下Kafka集群Broker下线后无法选举新Leader排查

Kafka Kraft集群故障恢复问题

环境配置

集群架构

3个独立控制器 + 3个Broker,均采用Kraft模式,基于官方config/kraft下的示例配置迁移至config/server.properties。

控制器配置(每个控制器node.id分别为1/2/3)

node.id=1 # Each controller with its own id (1 to 3)
controller.quorum.voters=1@localhost:9193,2@localhost:9293,3@localhost:9393
listeners=CONTROLLER://:9393
num.partitions=10
num.recovery.threads.per.data.dir=2
offsets.topic.replication.factor=2
transaction.state.log.replication.factor=2
transaction.state.log.min.isr=2

Broker配置(每个Broker node.id分别为4/5/6)

node.id=4 # Each broker with its own id (4 to 6)
controller.quorum.voters=1@localhost:9193,2@localhost:9293,3@localhost:9393
listeners=PLAINTEXT://localhost:9492
inter.broker.listener.name=PLAINTEXT
advertised.listeners=PLAINTEXT://localhost:9492
controller.listener.names=CONTROLLER
num.partitions=10
num.recovery.threads.per.data.dir=2
offsets.topic.replication.factor=2
transaction.state.log.replication.factor=2
transaction.state.log.min.isr=2

启动脚本

每个节点执行:

bin/kafka-storage.sh format -t ffQlsx4mQn-ipKduywm2Ig -c config/server.properties --ignore-formatted
bin/kafka-server-start.sh config/server.properties

问题现象

初始集群运行正常,但关闭任意Broker后,集群完全无法工作:

  • 客户端收到Connection error: connect ECONNREFUSED 127.0.0.1:9492
  • 大量There is no leader for this topic-partition as we are in the middle of a leadership election错误,等待10分钟仍无法完成Leader选举

客户端用KafkaJS持续发送消息的代码:

setInterval(async () => {
  await producer.send({
    topic: 'test-topic',
    messages: [ { value: 'text' } ],
  })
}, 500);

Broker下线时的错误日志:

{
  namespace: 'Producer',
  label: 'ERROR',
  log: {
    timestamp: '2023-09-26T16:35:56.593Z',
    message: 'Failed to send messages: This server is not the leader for that topic-partition',
    broker: undefined,
    clientId: undefined,
    error: undefined,
    logLevel: 1
  }
}

几秒后错误变为:

{
  namespace: 'Connection',
  label: 'ERROR',
  log: {
    timestamp: '2023-09-26T16:36:05.807Z',
    message: 'Response Metadata(key: 3, version: 6)',
    broker: 'localhost:9692',
    clientId: '27a1ab62-7bd5-42e9-9c73-debfd284ac02',
    error: 'There is no leader for this topic-partition as we are in the middle of a leadership election',
    logLevel: 1
  }
}

以及:

{
  namespace: 'Producer',
  label: 'ERROR',
  log: {
    timestamp: '2023-09-26T16:36:06.098Z',
    message: 'Failed to send messages: Connection error: connect ECONNREFUSED 127.0.0.1:9492',
    broker: undefined,
    clientId: undefined,
    error: undefined,
    logLevel: 1
  }
}

疑问解答

1. 在Leader选举期间,是否需要手动处理producer.send操作?

需要手动处理。KafkaJS默认不会自动重试所有类型的错误(比如Leader选举中的无Leader错误),你需要在代码中添加重试逻辑:

  • 用try/catch包裹producer.send,捕获错误后延迟重试
  • 配置KafkaJS producer的重试参数,比如retries、retryBackoffMs,覆盖默认策略,确保对Leader选举类错误进行重试

示例代码片段:

const sendMessage = async () => {
  try {
    await producer.send({
      topic: 'test-topic',
      messages: [ { value: 'text' } ],
    });
  } catch (err) {
    console.error('发送失败,准备重试:', err.message);
    setTimeout(sendMessage, 1000);
  }
};
setInterval(sendMessage, 500);

2. Kafka本应具备故障恢复能力,是否我在环境配置中遗漏了某些设置?

你的配置存在几个关键问题,直接导致Leader选举失败:

  • 控制器端口不匹配:控制器的listeners配置为CONTROLLER://:9393,但controller.quorum.voters中对应的节点端口是9193/9293/9393,端口不一致导致控制器集群无法正常通信,Broker下线后无法触发Leader选举。需要将控制器的listeners端口与controller.quorum.voters中的对应端口对齐,比如node.id=1的控制器设置listeners=CONTROLLER://localhost:9193。
  • Broker端口重复:所有Broker的listeners和advertised.listeners都设置为localhost:9492,会导致端口冲突,实际只有一个Broker能成功启动。需要给每个Broker分配唯一端口,比如node.id=4用9492,node.id=5用9592,node.id=6用9692,同时更新advertised.listeners。
  • 副本因子配置不合理:offsets.topic.replication.factor=2,但你有3个Broker,建议设置为3,确保即使一个Broker下线,offsets主题仍有足够副本维持可用性;transaction.state.log.replication.factor同理也应设为3,transaction.state.log.min.isr可设为2。

3. 所有节点都在本地运行是否会影响测试结果的准确性?

本地运行本身不会影响测试准确性,但需要注意:

  • 确保每个节点使用独立的端口、数据目录,避免资源冲突(比如端口重复、数据目录重叠)
  • 本地机器的资源限制(CPU、内存)可能会影响集群的故障恢复速度,但不会导致完全无法选举Leader,你的问题核心还是配置错误

4. 在NodeJS应用与集群之间添加负载均衡器是否有帮助,还是这只是配置错误的临时解决方案?

负载均衡器无法解决当前的核心问题——集群无法正常进行Leader选举。它的作用是在生产环境中分发客户端请求到多个Broker,避免单点压力,但不能修复控制器通信失败、Broker端口冲突这类配置问题。你需要先解决配置错误,让集群具备正常的故障恢复能力,再考虑是否添加负载均衡器优化客户端访问。


内容的提问来源于stack exchange,提问作者Caio Favero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 00:47:05