如何在Elasticsearch中按唯一标签集合聚合查询桶?
按唯一标签集合聚合的Elasticsearch实现方案(无需新增字段)
现有索引与数据
先看已创建的索引结构和批量插入的数据:
创建索引
PUT /example { "mappings": { "properties": { "tags": { "type": "keyword" } } } }
批量插入数据
POST example/_bulk { "create" : { "_index" : "example" } } { "tags" : ["a", "b"] } { "create" : { "_index" : "example" } } { "tags" : ["c", "d"] } { "create" : { "_index" : "example" } } { "tags" : ["e"] } { "create" : { "_index" : "example" } } { "tags" : ["c", "d"] }
需求描述
需要对tags字段按唯一标签集合聚合,不是统计单个标签的文档数。比如示例里["c", "d"]出现两次,聚合后要显示该集合对应的文档数为2,最终结果结构要和以下示例一致:
{ ... "aggregations" : { "tags" : { "doc_count_error_upper_bound" : 0, "sum_other_doc_count" : 0, "buckets" : [ { "key" : ["a", "b"], "key_as_string" : "a|b", "doc_count" : 1 }, { "key" : ["c", "d"], "key_as_string" : "c|d", "doc_count" : 2 }, { "key" : ["e"], "key_as_string" : "e", "doc_count" : 1 } ] } } }
已知可以新增排序后的标签字符串字段来实现,但实际场景涉及嵌套字段,不想新增字段,所以寻求其他方案。
无需新增字段的实现方案
可以用Elasticsearch的脚本化terms聚合,或者结合runtime字段,通过脚本对tags数组排序后拼接成唯一字符串作为聚合key,就能实现按唯一标签集合聚合的效果。
方案一:脚本化terms聚合
直接在聚合里写脚本生成排序后的标签字符串,以此分组:
GET /example/_search { "size": 0, "aggs": { "unique_tag_sets": { "terms": { "script": { "source": """ // 对tags数组排序后拼接,确保顺序不同的同集合视为同一key def sortedTags = params._source.tags.stream().sorted().collect(Collectors.toList()); return sortedTags.join('|'); """ }, "size": 10 }, // 可选:添加子聚合获取数组格式的标签集合 "aggs": { "tag_set": { "top_hits": { "size": 1, "_source": ["tags"] } } } } } }
说明
- 脚本会把每个文档的
tags数组排序后用|拼接,比如["d", "c"]和["c", "d"]都会生成c|d,被分到同一个聚合桶里。 - 返回的
key是拼接后的字符串,doc_count就是该标签集合的文档数。如果需要数组格式的key,通过top_hits子聚合取第一个文档的tags即可,同一桶内的集合排序后是一致的。
方案二:基于runtime字段的聚合
先在查询时临时生成一个runtime字段,存储排序后的标签字符串,再基于这个字段聚合,写法更清晰:
GET /example/_search { "size": 0, "runtime_mappings": { "sorted_tags_str": { "type": "keyword", "script": """ def sortedTags = doc['tags'].stream().sorted().collect(Collectors.toList()); emit(sortedTags.join('|')); """ } }, "aggs": { "unique_tag_sets": { "terms": { "field": "sorted_tags_str", "size": 10 }, "aggs": { "tag_set": { "top_hits": { "size": 1, "_source": ["tags"] } } } } } }
嵌套字段场景适配
如果是嵌套字段(比如nested_tags.tags),只需结合nested聚合调整脚本里的字段路径:
GET /example/_search { "size": 0, "aggs": { "nested_tags": { "nested": { "path": "nested_tags" }, "aggs": { "unique_tag_sets": { "terms": { "script": """ def sortedTags = doc['nested_tags.tags'].stream().sorted().collect(Collectors.toList()); return sortedTags.join('|'); """ }, "aggs": { "tag_set": { "top_hits": { "size": 1, "_source": ["nested_tags.tags"] } } } } } } } }
内容的提问来源于stack exchange,提问作者Chris Starling
相关产品推荐
相关产品推荐

