如何优化Elasticsearch中含17万+文档索引的Percolator查询?
Hey there! Let's break down some practical optimizations for your Percolator query—especially since you're working with an index of ~170k documents. First, let's fix up your current query (I noticed a missing closing bracket to make it valid):
curl -XGET 'localhost:9200/title-index/_search?pretty' -H 'Content-Type: application/json' -d' { "query" : { "percolate" : { "field" : "query", "document" : { "title" : "Dışa, bireylerin bir ar..." } } } }'
Now, here are actionable tweaks to speed up your percolation performance:
1. Pre-filter Percolator Queries to Cut Down Scope
Percolator checks your input document against every stored query by default. If you can tag your percolator queries with categories (like topic, language, or department), add a filter to only match relevant queries. For example:
{ "query": { "bool": { "filter": [ { "term": { "content_topic": "external_relations" } } ], "must": [ { "percolate": { "field": "query", "document": { "title": "Dışa, bireylerin bir ar..." } } } ] } } }
This drastically reduces the number of queries your input document needs to be tested against.
2. Use Routing to Target Specific Shards
By default, Percolator scans all index shards. If you assign a consistent routing value (e.g., language code, category ID) to both your percolator queries and input documents, you can limit the query to only relevant shards. Example:
curl -XGET 'localhost:9200/title-index/_search?pretty&routing=turkish' -H 'Content-Type: application/json' -d' { "query" : { "percolate" : { "field" : "query", "document" : { "title" : "Dışa, bireylerin bir ar..." } } } }'
Just make sure you use the same routing value when indexing your percolator queries.
3. Optimize the Percolator Field Mapping
Ensure your query field is properly mapped as a percolator type, and add keyword indexes to any auxiliary fields you use for filtering. This makes pre-filtering faster:
{ "mappings": { "properties": { "query": { "type": "percolator" }, "content_topic": { "type": "keyword" } // Fast term filtering } } }
4. Batch Documents with the mpercolate API
If you need to match multiple input documents, don't send individual requests. Use the mpercolate API to batch them—this cuts down network overhead and lets Elasticsearch process them more efficiently:
curl -XGET 'localhost:9200/title-index/_mpercolate?pretty' -H 'Content-Type: application/json' -d' { "docs": [ { "percolate": { "field": "query", "document": { "title": "First document to match..." } } }, { "percolate": { "field": "query", "document": { "title": "Dışa, bireylerin bir ar..." } } } ] }'
5. Tune Shard and Replica Counts
For an index of 170k documents, check if your shard count is optimal. Aim for shards between 10-50GB each—too many small shards add overhead, too few large shards limit parallelism. Adjust replica counts too: more replicas boost read performance but increase write overhead, so balance based on your workload.
6. Enable Request Caching
If you run the same (or similar) Percolator queries repeatedly, enable Elasticsearch's request cache to store results. You can turn it on per query:
{ "query" : { "percolate" : { "field" : "query", "document" : { "title" : "Dışa, bireylerin bir ar..." } } }, "request_cache": true }
Or set it as a default in your index settings for repeated queries.
7. Simplify the Input Document
Only include fields your percolator queries actually use in the document section. If all queries target the title field, don't send extra fields—this reduces the data Elasticsearch needs to process during matching.
内容的提问来源于stack exchange,提问作者SalahAdDin

