Kraft模式下Kafka集群Broker下线后无法选举新Leader排查
Kafka Kraft集群故障恢复问题
环境配置
集群架构
3个独立控制器 + 3个Broker,均采用Kraft模式,基于官方config/kraft下的示例配置迁移至config/server.properties。
控制器配置(每个控制器node.id分别为1/2/3)
node.id=1 # Each controller with its own id (1 to 3) controller.quorum.voters=1@localhost:9193,2@localhost:9293,3@localhost:9393 listeners=CONTROLLER://:9393 num.partitions=10 num.recovery.threads.per.data.dir=2 offsets.topic.replication.factor=2 transaction.state.log.replication.factor=2 transaction.state.log.min.isr=2
Broker配置(每个Broker node.id分别为4/5/6)
node.id=4 # Each broker with its own id (4 to 6) controller.quorum.voters=1@localhost:9193,2@localhost:9293,3@localhost:9393 listeners=PLAINTEXT://localhost:9492 inter.broker.listener.name=PLAINTEXT advertised.listeners=PLAINTEXT://localhost:9492 controller.listener.names=CONTROLLER num.partitions=10 num.recovery.threads.per.data.dir=2 offsets.topic.replication.factor=2 transaction.state.log.replication.factor=2 transaction.state.log.min.isr=2
启动脚本
每个节点执行:
bin/kafka-storage.sh format -t ffQlsx4mQn-ipKduywm2Ig -c config/server.properties --ignore-formatted bin/kafka-server-start.sh config/server.properties
问题现象
初始集群运行正常,但关闭任意Broker后,集群完全无法工作:
- 客户端收到
Connection error: connect ECONNREFUSED 127.0.0.1:9492 - 大量
There is no leader for this topic-partition as we are in the middle of a leadership election错误,等待10分钟仍无法完成Leader选举
客户端用KafkaJS持续发送消息的代码:
setInterval(async () => { await producer.send({ topic: 'test-topic', messages: [ { value: 'text' } ], }) }, 500);
Broker下线时的错误日志:
{ namespace: 'Producer', label: 'ERROR', log: { timestamp: '2023-09-26T16:35:56.593Z', message: 'Failed to send messages: This server is not the leader for that topic-partition', broker: undefined, clientId: undefined, error: undefined, logLevel: 1 } }
几秒后错误变为:
{ namespace: 'Connection', label: 'ERROR', log: { timestamp: '2023-09-26T16:36:05.807Z', message: 'Response Metadata(key: 3, version: 6)', broker: 'localhost:9692', clientId: '27a1ab62-7bd5-42e9-9c73-debfd284ac02', error: 'There is no leader for this topic-partition as we are in the middle of a leadership election', logLevel: 1 } }
以及:
{ namespace: 'Producer', label: 'ERROR', log: { timestamp: '2023-09-26T16:36:06.098Z', message: 'Failed to send messages: Connection error: connect ECONNREFUSED 127.0.0.1:9492', broker: undefined, clientId: undefined, error: undefined, logLevel: 1 } }
疑问解答
1. 在Leader选举期间,是否需要手动处理producer.send操作?
需要手动处理。KafkaJS默认不会自动重试所有类型的错误(比如Leader选举中的无Leader错误),你需要在代码中添加重试逻辑:
- 用
try/catch包裹producer.send,捕获错误后延迟重试 - 配置KafkaJS producer的重试参数,比如
retries、retryBackoffMs,覆盖默认策略,确保对Leader选举类错误进行重试
示例代码片段:
const sendMessage = async () => { try { await producer.send({ topic: 'test-topic', messages: [ { value: 'text' } ], }); } catch (err) { console.error('发送失败,准备重试:', err.message); setTimeout(sendMessage, 1000); } }; setInterval(sendMessage, 500);
2. Kafka本应具备故障恢复能力,是否我在环境配置中遗漏了某些设置?
你的配置存在几个关键问题,直接导致Leader选举失败:
- 控制器端口不匹配:控制器的
listeners配置为CONTROLLER://:9393,但controller.quorum.voters中对应的节点端口是9193/9293/9393,端口不一致导致控制器集群无法正常通信,Broker下线后无法触发Leader选举。需要将控制器的listeners端口与controller.quorum.voters中的对应端口对齐,比如node.id=1的控制器设置listeners=CONTROLLER://localhost:9193。 - Broker端口重复:所有Broker的
listeners和advertised.listeners都设置为localhost:9492,会导致端口冲突,实际只有一个Broker能成功启动。需要给每个Broker分配唯一端口,比如node.id=4用9492,node.id=5用9592,node.id=6用9692,同时更新advertised.listeners。 - 副本因子配置不合理:
offsets.topic.replication.factor=2,但你有3个Broker,建议设置为3,确保即使一个Broker下线,offsets主题仍有足够副本维持可用性;transaction.state.log.replication.factor同理也应设为3,transaction.state.log.min.isr可设为2。
3. 所有节点都在本地运行是否会影响测试结果的准确性?
本地运行本身不会影响测试准确性,但需要注意:
- 确保每个节点使用独立的端口、数据目录,避免资源冲突(比如端口重复、数据目录重叠)
- 本地机器的资源限制(CPU、内存)可能会影响集群的故障恢复速度,但不会导致完全无法选举Leader,你的问题核心还是配置错误
4. 在NodeJS应用与集群之间添加负载均衡器是否有帮助,还是这只是配置错误的临时解决方案?
负载均衡器无法解决当前的核心问题——集群无法正常进行Leader选举。它的作用是在生产环境中分发客户端请求到多个Broker,避免单点压力,但不能修复控制器通信失败、Broker端口冲突这类配置问题。你需要先解决配置错误,让集群具备正常的故障恢复能力,再考虑是否添加负载均衡器优化客户端访问。
内容的提问来源于stack exchange,提问作者Caio Favero
相关产品推荐
相关产品推荐

