Elasticsearch聚合查询过慢甚至超时问题求助
Hey there! Let's figure out why your aggregation queries are running super slow or timing out, and walk through fixes that’ll get things moving again. First, let's recap your setup to ground our fixes:
ES INFO: nodes: 1, primary-shards: 30, replica-shards: 30, total index size: 90gb, each shards holding 3gb of data, SYSTEM INFO: 250 SSD ram 28GB
In a single-node cluster, replica shards do absolutely nothing — Elasticsearch won’t assign replicas to the same node (they’re meant for redundancy across nodes). Those 30 replicas are just cluttering up cluster metadata and wasting resources that could go towards your queries. Fix this immediately by setting replica count to 0:
PUT /_all/_settings { "number_of_replicas": 0 }
This will free up overhead right away and let your node focus on the primary shards that actually hold data.
30 primary shards on a single node is way more than you need. As a rule of thumb, single-node clusters should stick to 1-2 primary shards per CPU core (if your server has 8 cores, 8-16 shards is ideal). Too many shards force the node to waste extra resources coordinating between them, merging segments, and managing metadata.
You can’t change shard count on an existing index, so you’ll need to reindex:
- Create a new index with a reasonable primary shard count (e.g., 10) and
number_of_replicas: 0 - Use the
_reindexAPI to copy data from your old index to the new one - Switch over to using the new index (or update your alias if you’re using one)
Elasticsearch relies heavily on its heap memory and the OS file cache. With 28GB of physical RAM, set your heap to 14GB (half the physical RAM) — leave the other 14GB for the OS file cache, which will speed up data reads from your SSD dramatically.
Edit your config/jvm.options file:
-Xms14g -Xmx14g
Important: Never set heap above 32GB — this disables JVM compressed pointers and wastes memory.
Sometimes the issue isn’t the cluster setup, but the query itself. Try these tweaks:
- Filter first, aggregate later: Use a
filterclause to narrow down to only the data you need before running the aggregation (e.g., if you only care about the last 30 days, add arangefilter on your timestamp field). This reduces the total data the aggregation has to process. - Balance precision and speed for high-cardinality fields: If you’re using a
cardinalityaggregation, adjust theprecision_thresholdparameter to trade a tiny bit of accuracy for faster results. - Avoid scripts in aggregations: Scripts add massive overhead — rewrite your query to use built-in aggregation functions instead.
- Use approximate aggregations if possible: If you don’t need exact results, try
samplerordiversified_sampleraggregations to run the aggregation on a sample of your data, which cuts down runtime significantly.
- Check SSD IO bottlenecks: Use tools like
iostatortopto see if your disk is maxed out. If it is, adjust the segment merge policy: setindex.merge.policy.max_merged_segmentto a larger value (e.g., 10gb) to reduce how often ES merges index segments. - Reduce refresh overhead: If you don’t need real-time search, set
index.refresh_intervalto 30s or longer. This reduces how often ES makes new data searchable, saving CPU and IO. - Increase file descriptors: Elasticsearch needs lots of file descriptors — set your system’s limit to at least 65535 to avoid errors and slowdowns.
With these changes, you should see a huge improvement in your aggregation query speeds. Start with the replica shard and heap memory fixes first — those are quick wins that’ll make an immediate difference.
内容的提问来源于stack exchange,提问作者ARUN ARUMUGAM

