You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ElasticSearch带Term聚合查询过慢问题求助

关于Elasticsearch聚合性能异常的问题

我的Elasticsearch集群存储了约2000万条数据,仅执行以下带uniqueId条件的term查询时,耗时仅1ms:

"query": {
  "bool": {
    "must": [
      {
        "term": {
          "uniqueId": "10921152912079"
        }
      }
    ]
  }
}

但在该过滤查询基础上添加嵌套terms聚合(基于userId和randomId字段)后,查询耗时升至约2秒,查询语句如下:

"query": {
  "bool": {
    "must": [
      {
        "term": {
          "uniqueId": "10921152912079"
        }
      }
    ]
  }
},
"aggs": {
  "users": {
    "terms": {
      "field": "userId"
    },
    "aggs": {
      "raf": {
        "terms": {
          "field": "randomId"
        }
      }
    }
  }
}

疑问

明明过滤后的结果最多只有10条文档,为何聚合耗时如此之久?是否聚合未基于过滤结果执行?性能分析显示聚合阶段耗时极高,相关信息如下:

"aggregations" : [
  {
    "type" : "GlobalOrdinalsStringTermsAggregator",
    "description" : "users",
    "time_in_nanos" : 91737162,
    "breakdown" : {
      "reduce" : 0,
      "build_aggregation_count" : 1,
      "post_collection" : 1878,
      "initialize_count" : 1,
      "reduce_count" : 0,
      "collect_count" : 5,
      "post_collection_count" : 1,
      "build_leaf_collector" : 15302476,
      "build_aggregation" : 76398285,
      "build_leaf_collector_count" : 28,
      "initialize" : 23555,
      "collect" : 10968
    },
    "debug" : {
      "segments_with_multi_valued_ords" : 0,
      "collection_strategy" : "dense",
      "segments_with_single_valued_ords" : 28,
      "deferred_aggregators" : [
        "raf"
      ],
      "total_buckets" : 9942540,
      "built_buckets" : 1,
      "result_strategy" : "terms",
      "has_filter" : false
    }
  }
]

问题分析与解决方案

核心问题定位

从性能数据的debug字段可以看出两个关键异常:

  • "has_filter" : false:聚合未识别到查询的过滤条件,没有基于过滤后的小结果集执行
  • "total_buckets" : 9942540:聚合提前构建了近1000万个userId的全局桶,这是耗时飙升的根本原因

原因解释

Elasticsearch的terms聚合默认使用**全局序数(Global Ordinals)**优化性能,该机制会提前为字段构建全局唯一值映射。但当查询过滤后的结果集极小(仅10条)时,全局序数的构建开销(扫描所有28个分段、构建近千万个桶)远大于直接扫描匹配文档的开销,导致性能严重下降。

解决方案

  1. 强制聚合使用匹配文档扫描模式
    在terms聚合中添加"execution_hint": "map",让聚合仅处理查询匹配到的文档,跳过全局序数构建:

    "aggs": {
      "users": {
        "terms": {
          "field": "userId",
          "execution_hint": "map"
        },
        "aggs": {
          "raf": {
            "terms": {
              "field": "randomId",
              "execution_hint": "map"
            }
          }
        }
      }
    }
    
  2. 验证字段映射正确性
    确认userId和randomId字段为keyword类型(ID类字段不应分词),避免因分词导致全局序数的基数膨胀。

  3. 优化索引分段
    当前索引有28个分段,可在只读索引上执行_forcemerge操作合并分段,减少聚合时需要扫描的分段数量:

    PUT /your-index/_forcemerge?max_num_segments=5
    
  4. 确认过滤结果的准确性
    在查询中添加"size": 0, "track_total_hits": true,验证实际匹配的文档数是否符合预期,避免过滤条件未正确生效的情况:

    "query": { ... },
    "size": 0,
    "track_total_hits": true,
    "aggs": { ... }
    

内容的提问来源于stack exchange,提问作者user1590595

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 08:12:51