Elasticsearch中基于列表内字段的文档计数聚合问题
嘿,我来帮你搞定这个Elasticsearch的聚合问题!你的核心需求是统计每个作者对应的文档数,哪怕一篇文档有多个作者,每个作者都要单独算一次——这个需求其实不难实现,咱们分两种情况来解决:
快速搞定(不用改索引映射)
你的authors字段默认是ES的object类型,这种类型会把数组里的对象扁平化存储。你之前踩的坑是用了authors.keyword(这是把整个作者对象当成关键字),而不是针对name字段来聚合。直接对authors.name.keyword做terms聚合就能得到你要的结果:
curl -X GET "localhost:9200/news/_search?pretty" -H 'Content-Type: application/json' -d' { "size": 0, # 不需要返回具体文档,只看聚合结果 "aggs": { "author_count": { "terms": { "field": "authors.name.keyword", "size": 100 # 按需调整显示的作者数量,默认只返回前10个 } } } }
如果这个请求还是没返回任何桶,先检查下你的索引映射,确认authors.name字段有没有keyword子字段。用下面的命令看映射:
curl -X GET "localhost:9200/news/_mapping?pretty"
要是name只有text类型没有keyword,可以用脚本聚合来提取名字:
curl -X GET "localhost:9200/news/_search?pretty" -H 'Content-Type: application/json' -d' { "size": 0, "aggs": { "author_count": { "terms": { "script": "doc[''authors.name''].values", "size": 100 } } } }
长期优化方案:改成Nested类型映射
如果之后你需要对作者的多个字段(比如name和@type)做更复杂的关联聚合,或者想避免object类型扁平化带来的潜在问题,建议把authors改成nested类型。步骤很简单:
- 先建一个带nested映射的新索引:
curl -X PUT "localhost:9200/news_new?pretty" -H 'Content-Type: application/json' -d' { "mappings": { "properties": { "authors": { "type": "nested", # 标记为nested类型 "properties": { "name": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } }, "@type": { "type": "keyword" } } }, "resort": { "type": "keyword" } # 其他原有字段都要在这里定义哦 } } }
- 把旧索引的数据迁移到新索引:
curl -X POST "localhost:9200/_reindex?pretty" -H 'Content-Type: application/json' -d' { "source": { "index": "news" }, "dest": { "index": "news_new" } }
- 用nested聚合统计作者数:
现在你之前尝试的聚合就能正常工作了,而且能精准处理每个独立的作者对象:
curl -X GET "localhost:9200/news_new/_search?pretty" -H 'Content-Type: application/json' -d' { "size": 0, "aggs": { "nested_authors": { "nested": { "path": "authors" }, "aggs": { "author_count": { "terms": { "field": "authors.name.keyword", "size": 100 } } } } } }
为啥你之前的nested聚合报错?
因为你的authors字段默认是object类型,不是nested类型,ES找不到对应的嵌套路径,自然就报错啦。只有当字段被明确设置为nested类型时,才能用nested聚合来处理数组里的对象。
内容的提问来源于stack exchange,提问作者chrisl
相关产品推荐
相关产品推荐

