You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

生产者ACK延迟过高时需监控的Kafka Broker指标(Datadog环境)

Key Broker-Side Lag & Load Metrics to Diagnose High Producer ACK Latency

When dealing with producer ACK latency spiking over 10 seconds, sticking only to throughput metrics like message.in.rate or kafka.net.bytes_in.rate misses the critical context of how Broker load and message backlogs are dragging down acknowledgment times. Here are the most impactful lag-related and associated load metrics you should monitor in Datadog to pinpoint the root cause:

Core Lag Metrics Tied to Broker Capacity

  • kafka.consumer.group.lag: This tracks the total unprocessed messages across all partitions for a consumer group. A sustained upward trend here means the Broker can’t keep pace with producer message volumes. As backlogs build on partitions, producers have to wait longer for their messages to be written, replicated, and acknowledged—directly pushing ACK latency higher.
  • kafka.partition.lag: Drill down to individual partition-level lag with this metric. If specific partitions show abnormally high lag, it’s a clear sign the Broker hosting those partitions is overloaded (e.g., saturated disk IO, maxed-out CPU). The backlog on these partitions slows new message processing, leading to delayed ACKs for producers targeting those partitions.
  • kafka.replica.lag: This measures the gap between a leader partition and its in-sync replicas (ISRs). For producers using acks=all (or -1), the Broker must wait for all ISRs to confirm receipt before sending an ACK. If replica lag spikes, follower replicas can’t keep up with the leader’s replication speed—often due to network bandwidth limits or slow disk writes on follower Brokers. This directly prolongs the ACK wait time.

Associated Broker Load Metrics That Amplify Lag

While not strictly "lag" metrics, these directly correlate with lag buildup and ACK latency:

  • kafka.log.flush.time.ms: The average time it takes the Broker to flush messages from memory to disk. A rise here points to disk performance degradation. For producers using acks=1 or acks=all (with flush requirements), the Broker can’t send an ACK until the message is safely written to disk—so slower flushes equal longer ACK delays.
  • kafka.network.request.wait.time.ms: This tracks how long producer requests sit in the Broker’s queue before being processed. If this metric climbs, the Broker’s CPU or memory is overloaded, and incoming requests are backing up. Producers will wait longer for their requests to be handled, resulting in higher ACK latency.
  • kafka.partition.under_replicated_partitions: A non-zero value here means some partitions have replicas out of sync with the leader. This forces the Broker to divert resources to replica synchronization, taking away from processing new producer messages. For producers expecting acks=all, the Broker may also take longer to confirm ACKs if fewer replicas are in sync.

How to Correlate These Metrics

Start by checking kafka.replica.lag and kafka.partition.lag—if either is spiking, cross-reference with the Broker’s resource metrics (e.g., disk.utilization, cpu.system, network.out.bytes):

  • High kafka.replica.lag + elevated disk.utilization on follower Brokers points to disk IO bottlenecks slowing replication.
  • High kafka.partition.lag + high kafka.network.request.wait.time.ms signals the Broker is overwhelmed with requests and can’t keep up with message processing.

内容的提问来源于stack exchange,提问作者Xiaohe Dong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:40:59