如何加速获取Elasticsearch多索引的唯一值?
First, let's fix critical issues in your current code that are wasting resources and failing to properly retrieve aggregation results:
Corrected Base Code
Your original code has two major flaws:
QueryBuilders.matchAllQueryis a method, so you need to invoke it withQueryBuilders.matchAllQuery()- You're not extracting aggregation results (fetching hits is irrelevant when using terms aggregations)
Here's the fixed code that properly captures unique values:
String[] instanceNames = getAllIndices().toArray(String[]::new); Map<String, Set<String>> indexToUniqueProviders = new HashMap<>(); // Query all indices in a single request (massive speedup over looping) SearchRequest searchRequest = new SearchRequest(instanceNames); SearchSourceBuilder searchSourceBuilder = new SearchSourceBuilder(); searchSourceBuilder.size(0); // Skip fetching hits—we only care about aggregations searchSourceBuilder.aggregation(AggregationBuilders.terms("DISTINCT_VALUES") .field("provider.keyword") .size(Integer.MAX_VALUE)); // Adjust if you have millions of unique values (use composite agg instead) searchRequest.source(searchSourceBuilder); try { SearchResponse searchResponse = getClient().search(searchRequest, RequestOptions.DEFAULT); Terms termsAgg = searchResponse.getAggregations().get("DISTINCT_VALUES"); Set<String> uniqueProviders = new HashSet<>(); for (Terms.Bucket bucket : termsAgg.getBuckets()) { uniqueProviders.add(bucket.getKeyAsString()); } // If you need per-index unique values, use multi-terms agg (see below) } catch (ElasticsearchException | IOException e) { throw new ServiceException(I18n.ELASTIC_SEARCH_ERROR, e); }
Key Optimization Strategies
1. Batch All Indices into One Request
Looping through each index and sending individual requests is the primary cause of slowness—each request incurs HTTP connection overhead. Send a single request targeting all relevant indices to eliminate this overhead entirely.
2. Disable Hit Retrieval
Add searchSourceBuilder.size(0) to skip fetching document hits. This reduces data transfer and processing time since you only need aggregation results.
3. Handle Large Unique Value Sets
- The default
termsaggregation returns only 10 buckets. Set.size()to a value that covers your expected unique count (e.g.,10000if you have up to 10k unique providers). - For datasets with 100k+ unique values, use the composite aggregation to paginate results without loading all data into memory at once:
CompositeAggregationBuilder compositeAgg = AggregationBuilders.composite("DISTINCT_VALUES") .sources(List.of(new TermsValuesSourceBuilder("provider").field("provider.keyword"))) .size(1000); // Adjust page size based on your cluster's capacity searchSourceBuilder.aggregation(compositeAgg);
Use the after parameter in subsequent requests to fetch additional pages.
4. Parallelize Batched Requests (If Needed)
If querying all indices at once causes timeouts, split indices into smaller batches (e.g., 10 indices per batch) and process them asynchronously with searchAsync(). Limit concurrent requests to 5-10 to avoid overwhelming your Elasticsearch cluster.
5. Filter Irrelevant Indices
Exclude indices that don't contain the provider.keyword field, or indices that are closed, archived, or unused. This reduces the amount of data Elasticsearch needs to scan.
6. Validate Field Configuration
Ensure provider.keyword has doc_values enabled (default for keyword fields). Doc_values use disk-based structures for faster aggregations instead of loading data into heap memory.
Get Per-Index Unique Values
If you need unique values broken down by index, use a multi-terms aggregation to group results by index and provider:
searchSourceBuilder.aggregation(AggregationBuilders.multiTerms("INDEX_PROVIDER") .addField("_index") .addField("provider.keyword") .size(Integer.MAX_VALUE));
Extract buckets to map each index to its unique provider values.
内容的提问来源于stack exchange,提问作者crosbee

