如何优化高写入负载下的Elasticsearch查询以降低延迟?
Elasticsearch查询性能优化问题
我的Elasticsearch查询多数情况下执行时间超过1分钟,被查询的索引每分钟至少接收500次写入更新。
使用的查询语句
GET /<my-index>/_search { "query": { "bool": { "must": [], "should": [], "filter": [ { "term": { "deleted": false } }, { "term": { "companyId.keyword": "my-companyId" } }, { "term": { "status.keyword": "booked" } }, { "terms": { "contactId.keyword": [ ... All ids (can provide upto 100 ids) ] } }, { "range": { "startTime": { "gte": <timestamp> } } } ], "must_not": [] } }, "size": <length of contactIds>, "from": 0, "aggs": { "groupByContactId": { "terms": { "field": "contactId.keyword", "size": <length of contactIds> }, "aggs": { "nextEvents": { "top_hits": { "size": 1, "sort": [ { "startTime": { "order": "asc" } } ] } } } } }, "sort": [ { "dateAdded": { "missing": -9999999999999, "order": "desc" } }, { "_id": { "order": "asc" } } ], "track_total_hits": false }
字段配置
companyId、contactId、status为text+keyword类型字段startTime、dateAdded为date类型字段- 所有字段均设置
index: true和doc_values: true
我的需求是为每个contactId查找一条即将发生的文档,但查询在部分时段性能不佳。
已尝试的优化措施
- 聚合中设置
execution_hint: map - 移除
dateAdded和_id的排序 - 在
top_hits中添加_source.includes选项
以上优化均未显著降低延迟。
搜索分析结果
GlobalOrdinalsStringTermsAggregator (groupByContactId) Time Spent: Ranges from ~672K ns to 1.35M ns across shards. Key Time Consumers: build_leaf_collector: Dominates with over 600K–1.4M ns. build_aggregation: Non-trivial (~40K–450K ns). ---------------------And in one shard Query had a high build_scorer time (~3.6M ns)
我尝试将旧索引重索引至新索引,并为contactId设置eager_global_ordinals: true,但整体查询延迟无显著下降。注:contactId为20字符长的字母数字ID。
优化建议
- 替换聚合方式,用
top_metrics替代terms+top_hits:top_metrics是专门针对「按分组取Top N文档」场景设计的聚合,性能远优于terms嵌套top_hits。它无需先构建全局序数再分组,直接在遍历文档时记录每个分组的Top结果,能大幅降低build_leaf_collector和聚合构建的开销。示例写法:"aggs": { "groupByContactId": { "top_metrics": { "metrics": [{"field": "startTime"}, {"field": "_source"}], "size": <length of contactIds>, "sort": [{"startTime": "asc"}], "group_by": {"terms": {"field": "contactId.keyword"}} } } } - 精简字段类型,减少索引开销:
contactId是字母数字ID,若不需要对其做全文搜索,可移除text类型,仅保留keyword类型。这样能减少索引存储容量和维护成本,降低全局序数构建的复杂度。 - 利用路由缩小查询范围:查询中固定包含
companyId过滤条件,可将companyId设为索引的路由键。创建索引时设置"routing": {"required": true},写入文档时指定routing为companyId,查询时同样携带routing=my-companyId参数,这样查询只会命中对应分片,避免遍历所有分片,大幅减少查询开销。 - 关闭无用的查询结果返回:你的需求是获取聚合结果,查询的
hits部分完全可以设置size: 0,减少数据传输和处理的资源消耗。 - 调整分片与刷新策略:
- 检查分片大小,将单个分片容量控制在20-50GB之间,避免分片过大导致查询压力集中;
- 适当调大
refresh_interval(比如从默认1s改为5s),减少写入时的segment生成频率,降低查询时遍历过多segment的开销。
- 优化全局序数加载策略:若
contactId基数极高,eager_global_ordinals可能因频繁更新导致开销过大,可尝试设置"eager_global_ordinals": {"loading": "eager_prime"},让节点在启动时预加载全局序数,减少运行时的构建开销。 - 检查集群资源瓶颈:确认节点CPU、内存、磁盘IO是否充足。聚合操作对CPU和内存消耗较大,若硬件资源不足,可考虑扩容节点或升级硬件(比如将机械硬盘换成SSD)。
内容的提问来源于stack exchange,提问作者suvodipMondal
相关产品推荐
相关产品推荐

