Kafka消费者组重平衡与Group Coordinator异常问题咨询
Alright, let's dig into your Kafka Connect + Cassandra sink issue. Those repeated "marking coordinator dead" / "discovered group coordinator" logs and unexpected connector restarts are classic signs of consumer group instability—especially with your single-broker Kafka 1.0.0 setup. Here's a breakdown of why this is happening and actionable fixes:
Core Root Causes
1. Single Broker = Single Point of Failure for Coordinator
Kafka’s consumer group coordinator runs on a broker, and with only one broker in your cluster, that’s your only coordinator. Even brief blips like network latency, JVM GC pauses, or resource bottlenecks on the broker will cause your Connect workers and sink consumers to drop their heartbeat connection. The coordinator then marks the consumer/worker as dead, triggering a rebalance—which is exactly what you’re seeing in the logs.
2. Kafka 1.0.0’s Known Coordinator Stability Gaps
Kafka 1.0.0 is an outdated release (over 5 years old at this point), and it had several unresolved issues with coordinator logic and rebalance handling. Later versions (2.0+) fixed a ton of these edge cases, especially around single-broker setups and heartbeat timeout handling.
3. Misconfigured Timeout Parameters
If your consumer or Connect worker timeout settings are too strict, even minor delays can trigger coordinator timeouts:
- The default
session.timeout.ms(30s) might be too short if your broker or network has occasional lag. - If your sink consumers are blocked waiting for slow Cassandra writes, they can’t send heartbeats to the coordinator in time, leading to false "dead" markings.
- Connect workers themselves are part of the
connect-clusterconsumer group—so their timeout settings also affect cluster stability.
4. Broker Resource Constraints
A single broker carries all cluster load (producer traffic, consumer coordination, storage I/O). If it’s starved for CPU, memory, or disk I/O:
- JVM full GCs can pause the broker long enough to drop heartbeats.
- Slow disk writes can delay coordinator responses to heartbeat requests.
- Network congestion between the broker and Connect nodes can interrupt heartbeat flows.
Fixes & Mitigations
1. Upgrade Kafka (Most Impactful Fix)
First and foremost, upgrade to a supported, stable Kafka version. Versions like 2.8.x (LTS) or 3.3.x (current LTS) have massive improvements to coordinator stability, rebalance logic, and error handling. If you can’t jump to a recent version, at least upgrade to 1.1.x—it patches many critical coordinator bugs in 1.0.0.
2. Tune Timeout Parameters
Adjust these settings in your sink consumer configs and Connect worker properties:
- Consumer/Sink Configs:
- Set
session.timeout.msto 60000 (60s) instead of the default 30000. - Set
heartbeat.interval.msto ~1/3 of the session timeout (e.g., 20000ms) to ensure frequent heartbeat checks. - Increase
consumer.max.poll.interval.msto 300000 (5 minutes) to give consumers time to process slow Cassandra writes without being marked dead.
- Set
- Connect Worker Properties:
- Apply the same
session.timeout.msandheartbeat.interval.msvalues to theconnect-clustergroup. - Increase
rebalance.timeout.msto 120000 (2 minutes) to give Connect enough time to complete rebalances without force-stopping connectors.
- Apply the same
3. Fix Broker Resource Bottlenecks
- Memory: Ensure the broker’s JVM heap is sized appropriately (e.g.,
KAFKA_HEAP_OPTS="-Xmx4G -Xms4G"for a moderate workload) to avoid frequent full GCs. Usejstat -gc <broker-pid>to monitor GC activity. - Disk: Check if the broker’s storage is slow (use
iostatoriotop). If disk I/O is high, switch to SSD storage or add more disk capacity. - Network: Verify there’s no packet loss or latency between Connect nodes and the broker using
pingormtr. Fix any network issues (e.g., faulty switches, overloaded routers) that might interrupt heartbeats.
4. Optimize Cassandra Sink Performance
Slow Cassandra writes can block your sink consumers, preventing them from sending heartbeats:
- Increase
batch.sizein the Cassandra sink config to reduce the number of individual writes to Cassandra. - Check your Cassandra cluster’s health: Ensure all nodes are up, there’s no pending compaction, and read/write latencies are within acceptable limits.
- If your Cassandra cluster can handle it, adjust
tasks.maxto match your topic count (you already have 10, which is correct for 10 single-partition topics).
内容的提问来源于stack exchange,提问作者el323

