You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新搭建Kafka 2.7集群ISR不同步问题排查与优化建议咨询

Kafka 2.7集群ISR不同步问题排查与优化方案

问题背景

在RHEL 7.9版本服务器上搭建5节点Apache Kafka 2.7全新集群,完成安装后发现部分分区的ISR(同步副本集)未处于完全同步状态。

副本不同步的可能原因

  • 慢副本:follower副本长期无法跟上leader写入速度,常见原因是follower节点存在I/O瓶颈,消息追加速率低于从leader拉取的速率
  • 卡住的副本:follower副本停止从leader拉取数据,可能由GC停顿、节点故障或宕机导致
  • 引导中副本:提升主题副本因子后,新follower副本在追上leader日志前处于不同步状态

由于是全新集群,怀疑问题与server.properties配置参数相关,以下是__consumer_offsets主题的部分分区ISR异常信息:

Topic:__consumer_offsets        PartitionCount:50       ReplicationFactor:3     Configs:segment.bytes=104857600,cleanup.policy=compact,compression.type=producer
        Topic: __consumer_offsets       Partition: 0    Leader: 1003    Replicas: 1003,1001,1002        Isr: 1003,1001,1002
        Topic: __consumer_offsets       Partition: 1    Leader: 1001    Replicas: 1001,1002,1003        Isr: 1001,1003,1002
        Topic: __consumer_offsets       Partition: 2    Leader: 1003    Replicas: 1002,1003,1001        Isr: 1003,1001
        Topic: __consumer_offsets       Partition: 3    Leader: 1003    Replicas: 1003,1002,1001        Isr: 1003,1001
        Topic: __consumer_offsets       Partition: 4    Leader: 1001    Replicas: 1001,1003,1002        Isr: 1001,1003
        Topic: __consumer_offsets       Partition: 5    Leader: 1001    Replicas: 1002,1001,1003        Isr: 1003,1001,1002
        Topic: __consumer_offsets       Partition: 6    Leader: 1003    Replicas: 1003,1001,1002        Isr: 1003,1001,1002
        Topic: __consumer_offsets       Partition: 7    Leader: 1001    Replicas: 1001,1002,1003        Isr: 1001,1003,1002
        Topic: __consumer_offsets       Partition: 8    Leader: 1003    Replicas: 1002,1003,1001        Isr: 1003,1001
        Topic: __consumer_offsets       Partition: 9    Leader: 1003    Replicas: 1003,1002,1001        Isr: 1003,1001
        Topic: __consumer_offsets       Partition: 10   Leader: 1001    Replicas: 1001,1003,1002        Isr: 1001,1003
        Topic: __consumer_offsets       Partition: 11   Leader: 1001    Replicas: 1002,1001,1003        Isr: 1003

当前server.properties配置

auto.create.topics.enable=false
auto.leader.rebalance.enable=true
background.threads=10
log.retention.bytes=-1
log.retention.hours=12
delete.topic.enable=true
leader.imbalance.check.interval.seconds=300
leader.imbalance.per.broker.percentage=10
log.dir=/var/kafka/kafka-data
log.flush.interval.messages=9223372036854775807
log.flush.interval.ms=1000
log.flush.offset.checkpoint.interval.ms=60000
log.flush.scheduler.interval.ms=9223372036854775807
log.flush.start.offset.checkpoint.interval.ms=60000
compression.type=producer
log.roll.jitter.hours=0
log.segment.bytes=1073741824
log.segment.delete.delay.ms=60000
message.max.bytes=1000012
min.insync.replicas=1
num.io.threads=8
num.network.threads=3
num.recovery.threads.per.data.dir=1
num.replica.fetchers=1
offset.metadata.max.bytes=4096
offsets.commit.required.acks=-1
offsets.commit.timeout.ms=5000
offsets.load.buffer.size=5242880
offsets.retention.check.interval.ms=600000
offsets.retention.minutes=10080
offsets.topic.compression.codec=0
offsets.topic.num.partitions=50
offsets.topic.replication.factor=3
offsets.topic.segment.bytes=104857600
queued.max.requests=500
quota.consumer.default=9223372036854775807
quota.producer.default=9223372036854775807
replica.fetch.min.bytes=1
replica.fetch.wait.max.ms=500
replica.high.watermark.checkpoint.interval.ms=5000
replica.lag.time.max.ms=10000
replica.socket.receive.buffer.bytes=65536
replica.socket.timeout.ms=30000
request.timeout.ms=30000
socket.receive.buffer.bytes=102400
socket.request.max.bytes=104857600
socket.send.buffer.bytes=102400
transaction.max.timeout.ms=900000
transaction.state.log.load.buffer.size=5242880
transaction.state.log.min.isr=2
transaction.state.log.num.partitions=50
transaction.state.log.replication.factor=3
transaction.state.log.segment.bytes=104857600
transactional.id.expiration.ms=604800000
unclean.leader.election.enable=false
zookeeper.connection.timeout.ms=600000
zookeeper.max.in.flight.requests=10
zookeeper.session.timeout.ms=600000
zookeeper.set.acl=false
broker.id.generation.enable=true
connections.max.idle.ms=600000
connections.max.reauth.ms=0
controlled.shutdown.enable=true
controlled.shutdown.max.retries=3
controlled.shutdown.retry.backoff.ms=5000
controller.socket.timeout.ms=30000
default.replication.factor=2
delegation.token.expiry.time.ms=86400000
delegation.token.max.lifetime.ms=604800000
delete.records.purgatory.purge.interval.requests=1
fetch.purgatory.purge.interval.requests=1000
group.initial.rebalance.delay.ms=3000
group.max.session.timeout.ms=1800000
group.max.size=2147483647
group.min.session.timeout.ms=6000
log.cleaner.backoff.ms=15000
log.cleaner.dedupe.buffer.size=134217728
log.cleaner.delete.retention.ms=86400000
log.cleaner.enable=true
log.cleaner.io.buffer.load.factor=0.9
log.cleaner.io.buffer.size=524288
log.cleaner.io.max.bytes.per.second=1.7976931348623157e308
log.cleaner.max.compaction.lag.ms=9223372036854775807
log.cleaner.min.cleanable.ratio=0.5
log.cleaner.min.compaction.lag.ms=0
log.cleaner.threads=1
log.cleanup.policy=delete
log.index.interval.bytes=4096
log.index.size.max.bytes=10485760
log.message.timestamp.difference.max.ms=9223372036854775807
log.message.timestamp.type=CreateTime
log.preallocate=false
log.retention.check.interval.ms=300000
max.connections=2147483647
max.connections.per.ip=2147483647
max.incremental.fetch.session.cache.slots=1000
num.partitions=1
producer.purgatory.purge.interval.requests=1000
queued.max.request.bytes=-1
replica.fetch.backoff.ms=1000
replica.fetch.max.bytes=1048576
replica.fetch.response.max.bytes=10485760
reserved.broker.max.id=1500
transaction.abort.timed.out.transaction.cleanup.interval.ms=60000
transaction.remove.expired.transaction.cleanup.interval.ms=3600000
zookeeper.sync.time.ms=2000
broker.rack=/default-rack

(注:原配置中log.2131234cleaner.backoff.ms为无效配置,已修正为log.cleaner.backoff.ms`)

待验证方向

  • 逐步重启各Kafka broker
  • 删除不同步副本的数据(rm -rf对应分区目录),让Kafka重新同步
  • 使用kafka-reassign-partitions工具调整副本分布
  • 等待一段时间观察ISR是否自动同步
  • 将replica.lag.time.max.ms调整为1天并重启broker

优化建议

  1. 优先排查配置问题

    • log.flush.interval.ms=1000:强制每秒刷盘会极大增加I/O负载,建议改为默认值-1(由操作系统管理刷盘),或调整为更大值(如30000),减少不必要的磁盘写入压力
    • num.replica.fetchers=1:副本拉取线程数不足,建议增加到3-5,提升follower拉取leader数据的能力
    • replica.fetch.max.bytes=1048576:单批次拉取数据量偏小,可调整为5242880(5MB),减少拉取请求次数
  2. 验证方向点评

    • 逐步重启broker:对全新集群来说,可尝试先重启ISR缺失的follower节点,看是否能恢复同步
    • 删除副本数据:操作前需确认该节点无其他重要数据,删除后Kafka会自动从leader重新同步日志,这是较为直接的修复方式,但需注意业务无写入中断风险
    • kafka-reassign-partitions:适合调整副本分布,均衡节点负载,但需提前规划好副本分配方案,避免操作失误
    • 等待自动同步:若为引导中副本,短时间内可能自动追上,但全新集群出现该问题概率低,不建议长时间等待
    • 调整replica.lag.time.max.ms:仅延长副本被踢出ISR的时间,无法从根本解决同步慢的问题,不建议作为优先方案
  3. 额外检查项

    • 检查节点磁盘I/O使用率:使用iostat或vmstat查看follower节点的磁盘读写负载,确认是否存在I/O瓶颈
    • 检查JVM GC日志:查看是否有长时间GC停顿导致副本拉取中断
    • 确认节点网络连通性:确保各broker间网络无丢包、延迟过高情况

内容的提问来源于stack exchange,提问作者jessica

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 01:54:13