You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Elasticsearch中按唯一标签集合聚合查询桶?

按唯一标签集合聚合的Elasticsearch实现方案(无需新增字段)

现有索引与数据

先看已创建的索引结构和批量插入的数据:

创建索引

PUT /example
{
  "mappings": {
    "properties": {
      "tags": {
        "type": "keyword"
      }
    }
  }
}

批量插入数据

POST example/_bulk
{ "create" : { "_index" : "example" } }
{ "tags" : ["a", "b"] }
{ "create" : { "_index" : "example" } }
{ "tags" : ["c", "d"] }
{ "create" : { "_index" : "example" } }
{ "tags" : ["e"] }
{ "create" : { "_index" : "example" } }
{ "tags" : ["c", "d"] }

需求描述

需要对tags字段按唯一标签集合聚合,不是统计单个标签的文档数。比如示例里["c", "d"]出现两次,聚合后要显示该集合对应的文档数为2,最终结果结构要和以下示例一致:

{
  ...
  "aggregations" : {
    "tags" : {
      "doc_count_error_upper_bound" : 0,
      "sum_other_doc_count" : 0,
      "buckets" : [                       
        {
          "key" : ["a", "b"],
          "key_as_string" : "a|b",
          "doc_count" : 1
        },
        {
          "key" : ["c", "d"],
          "key_as_string" : "c|d",
          "doc_count" : 2
        },
        {
          "key" : ["e"],
          "key_as_string" : "e",
          "doc_count" : 1
        }
      ]
    }
  }
}

已知可以新增排序后的标签字符串字段来实现,但实际场景涉及嵌套字段,不想新增字段,所以寻求其他方案。

无需新增字段的实现方案

可以用Elasticsearch的脚本化terms聚合,或者结合runtime字段,通过脚本对tags数组排序后拼接成唯一字符串作为聚合key,就能实现按唯一标签集合聚合的效果。

方案一:脚本化terms聚合

直接在聚合里写脚本生成排序后的标签字符串,以此分组:

GET /example/_search
{
  "size": 0,
  "aggs": {
    "unique_tag_sets": {
      "terms": {
        "script": {
          "source": """
            // 对tags数组排序后拼接,确保顺序不同的同集合视为同一key
            def sortedTags = params._source.tags.stream().sorted().collect(Collectors.toList());
            return sortedTags.join('|');
          """
        },
        "size": 10
      },
      // 可选:添加子聚合获取数组格式的标签集合
      "aggs": {
        "tag_set": {
          "top_hits": {
            "size": 1,
            "_source": ["tags"]
          }
        }
      }
    }
  }
}

说明

  • 脚本会把每个文档的tags数组排序后用|拼接,比如["d", "c"]和["c", "d"]都会生成c|d,被分到同一个聚合桶里。
  • 返回的key是拼接后的字符串,doc_count就是该标签集合的文档数。如果需要数组格式的key,通过top_hits子聚合取第一个文档的tags即可,同一桶内的集合排序后是一致的。

方案二:基于runtime字段的聚合

先在查询时临时生成一个runtime字段,存储排序后的标签字符串,再基于这个字段聚合,写法更清晰:

GET /example/_search
{
  "size": 0,
  "runtime_mappings": {
    "sorted_tags_str": {
      "type": "keyword",
      "script": """
        def sortedTags = doc['tags'].stream().sorted().collect(Collectors.toList());
        emit(sortedTags.join('|'));
      """
    }
  },
  "aggs": {
    "unique_tag_sets": {
      "terms": {
        "field": "sorted_tags_str",
        "size": 10
      },
      "aggs": {
        "tag_set": {
          "top_hits": {
            "size": 1,
            "_source": ["tags"]
          }
        }
      }
    }
  }
}

嵌套字段场景适配

如果是嵌套字段(比如nested_tags.tags),只需结合nested聚合调整脚本里的字段路径:

GET /example/_search
{
  "size": 0,
  "aggs": {
    "nested_tags": {
      "nested": {
        "path": "nested_tags"
      },
      "aggs": {
        "unique_tag_sets": {
          "terms": {
            "script": """
              def sortedTags = doc['nested_tags.tags'].stream().sorted().collect(Collectors.toList());
              return sortedTags.join('|');
            """
          },
          "aggs": {
            "tag_set": {
              "top_hits": {
                "size": 1,
                "_source": ["nested_tags.tags"]
              }
            }
          }
        }
      }
    }
  }
}

内容的提问来源于stack exchange,提问作者Chris Starling

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 19:02:21