You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速获取Elasticsearch多索引的唯一值?

Optimization Solutions for Fetching Unique Values Across Multiple Elasticsearch Indices

First, let's fix critical issues in your current code that are wasting resources and failing to properly retrieve aggregation results:

Corrected Base Code

Your original code has two major flaws:

  • QueryBuilders.matchAllQuery is a method, so you need to invoke it with QueryBuilders.matchAllQuery()
  • You're not extracting aggregation results (fetching hits is irrelevant when using terms aggregations)

Here's the fixed code that properly captures unique values:

String[] instanceNames = getAllIndices().toArray(String[]::new);
Map<String, Set<String>> indexToUniqueProviders = new HashMap<>();

// Query all indices in a single request (massive speedup over looping)
SearchRequest searchRequest = new SearchRequest(instanceNames);
SearchSourceBuilder searchSourceBuilder = new SearchSourceBuilder();
searchSourceBuilder.size(0); // Skip fetching hits—we only care about aggregations
searchSourceBuilder.aggregation(AggregationBuilders.terms("DISTINCT_VALUES")
        .field("provider.keyword")
        .size(Integer.MAX_VALUE)); // Adjust if you have millions of unique values (use composite agg instead)

searchRequest.source(searchSourceBuilder);
try {
    SearchResponse searchResponse = getClient().search(searchRequest, RequestOptions.DEFAULT);
    Terms termsAgg = searchResponse.getAggregations().get("DISTINCT_VALUES");
    Set<String> uniqueProviders = new HashSet<>();
    for (Terms.Bucket bucket : termsAgg.getBuckets()) {
        uniqueProviders.add(bucket.getKeyAsString());
    }
    // If you need per-index unique values, use multi-terms agg (see below)
} catch (ElasticsearchException | IOException e) {
    throw new ServiceException(I18n.ELASTIC_SEARCH_ERROR, e);
}

Key Optimization Strategies

1. Batch All Indices into One Request

Looping through each index and sending individual requests is the primary cause of slowness—each request incurs HTTP connection overhead. Send a single request targeting all relevant indices to eliminate this overhead entirely.

2. Disable Hit Retrieval

Add searchSourceBuilder.size(0) to skip fetching document hits. This reduces data transfer and processing time since you only need aggregation results.

3. Handle Large Unique Value Sets

  • The default terms aggregation returns only 10 buckets. Set .size() to a value that covers your expected unique count (e.g., 10000 if you have up to 10k unique providers).
  • For datasets with 100k+ unique values, use the composite aggregation to paginate results without loading all data into memory at once:
CompositeAggregationBuilder compositeAgg = AggregationBuilders.composite("DISTINCT_VALUES")
        .sources(List.of(new TermsValuesSourceBuilder("provider").field("provider.keyword")))
        .size(1000); // Adjust page size based on your cluster's capacity

searchSourceBuilder.aggregation(compositeAgg);

Use the after parameter in subsequent requests to fetch additional pages.

4. Parallelize Batched Requests (If Needed)

If querying all indices at once causes timeouts, split indices into smaller batches (e.g., 10 indices per batch) and process them asynchronously with searchAsync(). Limit concurrent requests to 5-10 to avoid overwhelming your Elasticsearch cluster.

5. Filter Irrelevant Indices

Exclude indices that don't contain the provider.keyword field, or indices that are closed, archived, or unused. This reduces the amount of data Elasticsearch needs to scan.

6. Validate Field Configuration

Ensure provider.keyword has doc_values enabled (default for keyword fields). Doc_values use disk-based structures for faster aggregations instead of loading data into heap memory.

Get Per-Index Unique Values

If you need unique values broken down by index, use a multi-terms aggregation to group results by index and provider:

searchSourceBuilder.aggregation(AggregationBuilders.multiTerms("INDEX_PROVIDER")
        .addField("_index")
        .addField("provider.keyword")
        .size(Integer.MAX_VALUE));

Extract buckets to map each index to its unique provider values.

内容的提问来源于stack exchange,提问作者crosbee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 09:50:34