如何解决AWS托管Elasticsearch的Request Entity Too Large问题(Java RestHighClient)
Got it, let's tackle this 413 error you're hitting. The message makes it crystal clear: your bulk request is exceeding the 100MB default limit that AWS-managed Elasticsearch enforces. Here's how you can resolve this directly using the RestHighLevelClient library:
1. Split Your BulkRequest into Smaller Batches
The most reliable fix is to break up your large bulk request into smaller chunks that stay under the 100MB threshold (I recommend leaving a small buffer, like targeting 90MB per batch to avoid edge cases).
You'll need to track the size of each document as you add them to the bulk request, then send the batch once you hit your size limit. Here's a concrete example using Jackson to calculate document sizes:
import com.fasterxml.jackson.databind.ObjectMapper; import org.elasticsearch.action.bulk.BulkRequest; import org.elasticsearch.action.index.IndexRequest; import org.elasticsearch.client.RestHighLevelClient; import java.io.IOException; import java.util.List; public class BulkIndexer { private static final ObjectMapper OBJECT_MAPPER = new ObjectMapper(); // Set a safe batch size (90MB = 94371840 bytes) private static final long MAX_BATCH_SIZE_BYTES = 94371840; public void bulkIndexDocuments(RestHighLevelClient client, List<Object> documents, String indexName) throws IOException { BulkRequest bulkRequest = new BulkRequest(); long currentBatchSize = 0; for (Object doc : documents) { // Convert document to JSON and calculate its size in bytes byte[] docBytes = OBJECT_MAPPER.writeValueAsBytes(doc); long docSize = docBytes.length; // If adding this doc would exceed the limit, send the current batch if (currentBatchSize + docSize > MAX_BATCH_SIZE_BYTES) { client.bulk(bulkRequest); // Reset for next batch bulkRequest = new BulkRequest(); currentBatchSize = 0; } // Add the document to the batch bulkRequest.add(new IndexRequest(indexName).source(docBytes)); currentBatchSize += docSize; } // Send any remaining documents in the last batch if (!bulkRequest.requests().isEmpty()) { client.bulk(bulkRequest); } } }
2. Adjust RestHighLevelClient's HTTP Client Configuration
While splitting batches is the core fix, you should also ensure your client is configured to handle the maximum request size you're sending. RestHighLevelClient uses Apache HttpClient under the hood, so you can customize its request limits:
import org.elasticsearch.client.RestClient; import org.elasticsearch.client.RestClientBuilder; import org.elasticsearch.client.RestHighLevelClient; import org.apache.http.client.config.RequestConfig; import org.apache.http.impl.client.HttpClientBuilder; import org.apache.http.HttpHost; public class ClientFactory { public static RestHighLevelClient createClient(String esEndpoint) { RestClientBuilder builder = RestClient.builder(new HttpHost(esEndpoint, 443, "https")); builder.setHttpClientConfigCallback(httpClientBuilder -> { // Set the maximum allowed request size to match your batch limit RequestConfig requestConfig = RequestConfig.custom() .setContentLength(MAX_BATCH_SIZE_BYTES) .build(); return httpClientBuilder.setDefaultRequestConfig(requestConfig); }); return new RestHighLevelClient(builder); } }
Bonus: Request a Limit Increase from AWS (If Needed)
If splitting batches isn't feasible for your use case, you can reach out to AWS Support to increase the maximum request size for your Elasticsearch domain. Keep in mind that this has performance implications, so only do this if smaller batches aren't working for you.
内容的提问来源于stack exchange,提问作者Rahul Ranjan

