Confluent ElasticsearchSinkConnector偶发ConnectionClosedException不可恢复故障求助
Let’s break down how to tackle this intermittent ConnectionClosedException that’s causing your connector tasks to fail. This error typically stems from network-level timeouts, connection pool misconfigurations, or Elasticsearch (ES) cluster-side connection limits. Here are actionable steps to resolve it:
1. Tune Connection and Timeout Settings
Your current connection.timeout.ms and read.timeout.ms are set to 30 seconds, which might be too tight for large batch requests or periods of high ES cluster load. Additionally, managing idle connections can prevent stale, already-closed connections from being reused:
- Increase
read.timeout.msto 60000 (60 seconds) to give longer-running bulk requests more time to complete. - Add
connection.max.idle.time.msset to 300000 (5 minutes) to automatically prune idle connections before they get closed by network devices (like load balancers) or ES itself. - Explicitly set
max.retriesto 10 (it defaults to 5 in most versions) to allow more retry attempts for transient connection issues.
Example updated config snippets:
read.timeout.ms: 60000 connection.max.idle.time.ms: 300000 max.retries: 10
2. Adjust Batch Size and Retry Backoff
A batch size of 1000 might be generating oversized payloads that take too long to process, raising the risk of connection timeouts. Try scaling back to reduce per-request load:
- Lower
batch.sizeto 500 (or 200 if your records are large) to make each bulk request smaller and faster to handle. - You could also reduce
retry.backoff.msfrom 60000 to 10000 (10 seconds) to retry failed connections more quickly—just make sure your ES cluster can handle the additional load.
3. Check Elasticsearch Cluster Connection Limits
Your 3-node ES 7.8.0 cluster has a default http.max_open_connections setting of 200. With 10 Kafka Connect tasks, each potentially opening multiple connections, you might be hitting this limit:
- Check current connection usage via the ES API:
curl -u elastic:my-password https://es-http.com:9200/_nodes/http?pretty - If connections are approaching the limit, increase
http.max_open_connectionsin each node’selasticsearch.ymlto 500 or higher, then restart the ES nodes.
4. Inspect Network/Load Balancer Idle Timeouts
Since you’re connecting to a DNS name (es-http.com), traffic is likely routed through a load balancer. Most LBs have an idle connection timeout (often 60 seconds by default) that closes inactive connections. If your connector has long gaps between batches, it might reuse a stale connection that’s already been closed by the LB:
- Confirm the LB’s idle timeout setting, then set
connection.max.idle.time.msin your connector config to less than that value (e.g., 45000ms if the LB uses a 60s timeout). - Enable TCP keepalives on the LB if available to keep connections alive during idle periods.
5. Upgrade the Elasticsearch Sink Connector Version
Older versions of the Confluent Elasticsearch Sink Connector had known bugs with connection pool management, especially when handling closed connections. Since your ES cluster is on 7.8.0, ensure you’re using a connector version compatible with ES 7.x (e.g., Confluent 5.5.x to 6.2.x). Upgrading to the latest compatible version can resolve underlying connection-handling flaws.
6. Enable Additional Debug Logging
Even with errors.log.enable enabled, debug logs can reveal more context about connection lifecycle events. Add these lines to your Kafka Connect worker’s log4j configuration:
<Logger name="io.confluent.connect.elasticsearch" level="DEBUG"/> <Logger name="org.apache.http" level="DEBUG"/>
This will log detailed HTTP connection activity, helping you pinpoint exactly when and why connections are being closed.
内容的提问来源于stack exchange,提问作者Oleg Kaplun

