You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Elasticsearch中按桶键进行值计数并获取Top10 ID

我来帮你搞定这个问题!你当前的聚合查询里,value_count统计的是每个桶内文档中所有history.some_id的数量,而不是只统计和桶key一致的那个ID的出现次数,这就是结果不符合预期的原因。

给你两种解决方案,按需选择:

方案1:Filter聚合+Value Count

先通过filter过滤出当前桶key对应的some_id,再统计它的出现次数,逻辑直观:

{
  "size": 0,
  "aggs": {
    "group_by_id": {
      "terms": {
        "field": "history.some_id",
        "size": 10  // 指定取前10个高频ID
      },
      "aggs": {
        "filter_current_id": {
          "filter": {
            "term": {
              "history.some_id": "{{key}}"  // 用桶的key精准过滤
            }
          },
          "aggs": {
            "total_occurrences": {
              "value_count": {
                "field": "history.some_id"
              }
            }
          }
        }
      }
    }
  }
}

方案2:Script+Sum聚合(性能更优)

直接用脚本计算每个文档中当前桶key的出现次数,再求和,避免额外的filter开销:

{
  "size": 0,
  "aggs": {
    "group_by_id": {
      "terms": {
        "field": "history.some_id",
        "size": 10,
        "order": {
          "total_occurrences": "desc"  // 按总出现次数降序排列,确保取前10
        }
      },
      "aggs": {
        "total_occurrences": {
          "sum": {
            "script": {
              "source": "doc['history.some_id'].stream().filter(v -> v == params.key).count()",
              "params": {
                "key": "{{key}}"
              }
            }
          }
        }
      }
    }
  }
}

关键说明

  • 两种方案返回的total_occurrences.value都是对应some_id在所有history数组中的总出现次数(包括同一文档内的重复出现)。
  • 如果你的需求是统计包含该some_id的文档数,其实terms聚合自带的doc_count字段就够用,不需要额外的子聚合。

内容的提问来源于stack exchange,提问作者nipeco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:05:12