如何避免RestHighLevelClient将Elasticsearch服务端IP/端口加入黑名单及优化高频查询配置
Let's tackle your problem head-on—first, we'll break down why the blacklisting is happening, then fix it, and finally assess if your current setup is fit for high-frequency query workloads.
Why Your Elasticsearch Node Is Getting Blacklisted
The RestHighLevelClient uses a built-in NodeBlacklistTracker that automatically adds nodes to a blacklist if they fail 3 consecutive requests (the default threshold). Once blacklisted, the client avoids sending requests to that node for 1 minute (default expiry time) before trying again.
With high-frequency queries, even temporary glitches (like network blips, ES node load spikes, or slow responses hitting your timeout limits) can trigger this threshold quickly, leading to the debug logs you're seeing.
Steps to Prevent Blacklisting
Here are actionable tweaks to reduce or eliminate blacklisting:
1. Increase Retry Attempts
By default, the client retries failed requests 3 times. Bumping this up gives the node more chances to recover before hitting the blacklist threshold:
builder.setHttpClientConfigCallback(configurer -> { // Retry up to 5 times instead of 3; second param enables retry on redirects configurer.setRetryHandler(new DefaultRetryHandler(5, true)); // Keep your existing connection pool, keep-alive, and IO thread config here return configurer; });
2. Adjust Blacklist Expiry & Failure Threshold
If you can't avoid occasional failures, make the client retry blacklisted nodes faster, or raise the failure count needed to trigger blacklisting:
- Shorten blacklist expiry: Set this system property before initializing the client to make nodes come back online faster (e.g., 30 seconds instead of 1 minute):
System.setProperty("es.rest.client.node_blacklist.expire_after_ms", "30000"); - Custom failure listener: For full control, override the default failure logic to only blacklist nodes after more failures:
builder.setFailureListener(new RestClient.FailureListener() { private final AtomicInteger failureCounter = new AtomicInteger(0); private static final int BLACKLIST_THRESHOLD = 5; @Override public void onFailure(Node node) { int currentFailures = failureCounter.incrementAndGet(); // Only trigger default blacklisting if we hit our custom threshold if (currentFailures >= BLACKLIST_THRESHOLD) { super.onFailure(node); failureCounter.set(0); // Reset after blacklisting } } @Override public void onNodeRecovery(Node node) { failureCounter.set(0); // Reset when node comes back online } });
3. Tune Timeouts to Reduce False Failures
Your current timeout values are set to minutes, which can lead to long-running requests tying up connections and triggering unnecessary failures. For high-frequency queries, switch to shorter, more reasonable timeouts (e.g., 10-30 seconds):
builder.setRequestConfigCallback(configurer -> { configurer.setSocketTimeout(30 * 1000); // 30s timeout for request execution configurer.setConnectTimeout(10 * 1000); // 10s timeout for connection establishment configurer.setConnectionRequestTimeout(5000); // 5s timeout waiting for a connection from the pool configurer.setContentCompressionEnabled(true); return configurer; });
Is Your Current Configuration Suitable for High-Frequency Queries?
Your setup has some solid foundations, but there are key optimizations to make it more robust:
What's Working Well
- Connection pool limits:
setMaxConnTotalandsetMaxConnPerRouteprevent overwhelming the ES node with too many concurrent connections. - Compression: Enabling request/response compression reduces network overhead, which is critical for high-frequency workloads.
- Keep-alive strategy: Reusing connections cuts down on TCP handshake latency.
Areas to Improve
- Remove unnecessary headers: You're sending CORS-related headers (
ACCESS_CONTROL_ALLOW_ORIGIN, etc.) from the client—these only need to be configured on the ES server side. Removing them reduces request size and overhead. - Optimize connection pool sizing: Your current
setMaxConnPerRoute(Math.max(30, maxConnection / 10))is a safe default, but for a single ES node, you can setsetMaxConnPerRouteequal tosetMaxConnTotal(e.g., 200 total connections all dedicated to the single node) to maximize throughput, as long as the ES node'shttp.max_open_connectionssetting (default 1000) allows it. - IO thread count: Limiting to 15 threads is conservative for modern multi-core servers. Instead, set it to match your CPU capacity:
int ioThreadCount = Runtime.getRuntime().availableProcessors() * 2; configurer.setDefaultIOReactorConfig(IOReactorConfig.custom().setIoThreadCount(ioThreadCount).build()); - Connection request timeout: Setting it to
0means the client will wait indefinitely for a connection from the pool. This can lead to request backlogs during peak load—set a reasonable timeout (like 5 seconds) instead.
Final Optimized Client Configuration
Here's a revised version of your bean incorporating all these changes:
@Bean public RestHighLevelClient getElasticSearchClient() { // Shorten blacklist expiry to 30 seconds System.setProperty("es.rest.client.node_blacklist.expire_after_ms", "30000"); String encoding = Base64.getEncoder().encodeToString((username + ":" + password).getBytes()); Header[] headers = { new BasicHeader(HttpHeaders.CONTENT_TYPE, "application/json;charset=utf-8"), new BasicHeader(HttpHeaders.ACCEPT, "application/json"), new BasicHeader(HttpHeaders.AUTHORIZATION, "basic " + encoding) // Removed unnecessary CORS headers }; RestClientBuilder builder = RestClient.builder(new HttpHost(elasticSearchIp, elasticSearchPort)) .setHttpClientConfigCallback(configurer -> { int maxConnTotal = 200; configurer.setMaxConnTotal(maxConnTotal); // For single node, set max per route equal to total connections configurer.setMaxConnPerRoute(maxConnTotal); // Increase retry attempts to 5 configurer.setRetryHandler(new DefaultRetryHandler(5, true)); configurer.setKeepAliveStrategy((httpResponse, httpContext) -> keepAlive * 60 * 1000); // IO threads based on CPU cores int ioThreadCount = Runtime.getRuntime().availableProcessors() * 2; configurer.setDefaultIOReactorConfig( IOReactorConfig.custom().setIoThreadCount(ioThreadCount).build()); return configurer; }) .setRequestConfigCallback(configurer -> { configurer.setSocketTimeout(30 * 1000); configurer.setConnectTimeout(10 * 1000); configurer.setConnectionRequestTimeout(5000); configurer.setContentCompressionEnabled(true); return configurer; }) .setCompressionEnabled(true) .setDefaultHeaders(headers); // Custom failure listener to raise blacklist threshold builder.setFailureListener(new RestClient.FailureListener() { private final AtomicInteger failureCounter = new AtomicInteger(0); private static final int BLACKLIST_THRESHOLD = 5; @Override public void onFailure(Node node) { int currentFailures = failureCounter.incrementAndGet(); if (currentFailures >= BLACKLIST_THRESHOLD) { super.onFailure(node); failureCounter.set(0); } } @Override public void onNodeRecovery(Node node) { failureCounter.set(0); } }); return new RestHighLevelClient(builder); }
内容的提问来源于stack exchange,提问作者Hatef Alipoor

