You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RestHighLevelClient遇Connection reset by peer问题排查求助

Connection reset by peer with RestHighLevelClient after idle periods in AWS VPC

Let’s break down why you’re hitting this error and walk through actionable fixes. The "Connection reset by peer" message almost always means a TCP connection your client thought was active got closed by the other end (or a middleman like your AWS Load Balancer) during idle periods. Here’s the breakdown:

Root Causes

  1. AWS Load Balancer Idle Timeout: AWS Application Load Balancers (ALBs) have a default idle timeout of 60 seconds; Network Load Balancers (NLBs) default to 350 seconds. If your service goes longer than this without sending requests, the LB tears down the idle connection—but your client’s connection pool still holds that stale connection. When you send the next request, it tries to use the dead connection, triggering the reset error.
  2. Unconfigured Client Connection Pool: Your current code uses the default RestHighLevelClient settings, which don’t include idle connection validation or cleanup. The underlying Apache HttpClient doesn’t automatically check if a connection is alive before using it, so it happily hands out stale connections after long idle periods.
  3. Elasticsearch HTTP Idle Settings: ES 7.5 has a default http.max_idle_time of 10 seconds, meaning ES will close HTTP connections that stay idle longer than that. If your client holds a connection idle beyond this window, ES closes it, and the client has no idea until it tries to send a request.

Solutions

Let’s fix this starting with client-side changes (the most impactful):

1. Configure Connection Pool Validation & Cleanup

Modify your ESClientWrapper to add connection pool checks and idle cleanup, which prevents stale connections from being used. Here’s the updated code:

import org.apache.http.HttpHost;
import org.apache.http.client.config.RequestConfig;
import org.apache.http.impl.client.HttpClientBuilder;
import org.apache.http.impl.conn.PoolingHttpClientConnectionManager;
import org.elasticsearch.client.RestClient;
import org.elasticsearch.client.RestHighLevelClient;

import java.io.FileInputStream;
import java.io.IOException;
import java.util.Properties;
import java.util.concurrent.TimeUnit;

public class ESClientWrapper {
    private RestHighLevelClient client;

    public ESClientWrapper() throws IOException {
        FileInputStream propertiesFile = new FileInputStream("/var/elastic.properties");
        Properties properties = new Properties();
        properties.load(propertiesFile);
        
        // Set up connection manager with idle cleanup rules
        PoolingHttpClientConnectionManager connectionManager = new PoolingHttpClientConnectionManager();
        // Validate connections after 5 seconds of inactivity to catch stale ones
        connectionManager.setValidateAfterInactivity(5000);
        // Max idle time for connections (shorter than LB/ES idle timeout to avoid resets)
        connectionManager.setMaxIdleTime(45, TimeUnit.SECONDS);
        // Adjust max connections based on your service's load needs
        connectionManager.setMaxTotal(20);
        connectionManager.setDefaultMaxPerRoute(10);

        // Configure request timeouts to avoid hanging requests
        RequestConfig requestConfig = RequestConfig.custom()
                .setConnectTimeout(5000) // Timeout to establish a connection
                .setSocketTimeout(30000) // Timeout for data transfer
                .setConnectionRequestTimeout(5000) // Timeout to get a connection from the pool
                .build();

        // Enable TCP keep-alive to prevent middlemen from closing idle connections prematurely
        HttpClientBuilder httpClientBuilder = HttpClientBuilder.create()
                .setConnectionManager(connectionManager)
                .setDefaultRequestConfig(requestConfig)
                .setKeepAliveStrategy((response, context) -> 30000); // Send keep-alive every 30 seconds

        RestClientBuilder builder = RestClient.builder(
                new HttpHost(properties.getProperty("host"), 
                        Integer.parseInt(properties.getProperty("port"))))
                .setHttpClientConfigCallback(httpClientBuilder -> httpClientBuilder);

        this.client = new RestHighLevelClient(builder);
    }

    public RestHighLevelClient getClient() {
        return client;
    }
}

2. Adjust AWS Load Balancer Idle Timeout

If you’re using an ALB:

  • Go to the AWS Console → EC2 → Load Balancers → Select your LB → Attributes → Idle timeout.
  • Set it to a value longer than your client’s setMaxIdleTime (e.g., 60 seconds) but not excessively long (to avoid holding unnecessary connections).

For NLBs, the default 350-second timeout is usually acceptable if your client’s idle cleanup is configured correctly.

3. Add Retry Logic for Transient Errors

Wrap your ES calls in a retry loop to handle connection reset errors gracefully (transient issues like this often resolve with a retry):

public <T> T executeWithRetry(Supplier<T> esCall) throws IOException {
    int maxRetries = 3;
    int retryCount = 0;
    while (retryCount < maxRetries) {
        try {
            return esCall.get();
        } catch (IOException e) {
            if (e.getMessage().contains("Connection reset by peer") && retryCount < maxRetries - 1) {
                retryCount++;
                try {
                    TimeUnit.SECONDS.sleep(1);
                } catch (InterruptedException ie) {
                    Thread.currentThread().interrupt();
                    throw new IOException("Retry interrupted", ie);
                }
            } else {
                throw e;
            }
        }
    }
    throw new IOException("Max retries reached for ES request");
}

4. Verify Elasticsearch HTTP Settings

Check your ES cluster’s elasticsearch.yml for http.max_idle_time. If you want to keep connections alive longer, you can increase it (e.g., http.max_idle_time: 60s), but this is less critical if your client is already configured to validate and clean up idle connections.

Why It Works After 3 Minutes?

After the initial error, the client’s connection pool discards the stale connection and creates a new one. Alternatively, the connection manager’s idle cleanup kicks in after a few minutes, removing dead connections so subsequent requests get fresh, valid connections.

内容的提问来源于stack exchange,提问作者JeyJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 08:12:36