如何在Elasticsearch聚合中获取各Source的文档占比?
实现Instagram/Twitter文档占比的方案
完全可以实现需求,有两种常用方式:
方式一:在Elasticsearch聚合中直接计算占比
利用Elasticsearch的管道聚合(bucket_script),结合总文档数来计算每个来源的占比。修改后的聚合代码如下:
{ "aggs": { "total_docs": { "value_count": { "field": "_id" } }, "bySource": { "terms": { "field": "profile.source.keyword", "include": ["instagram", "twitter"] // 只统计指定来源,减少不必要计算 }, "aggs": { "source_percentage": { "bucket_script": { "buckets_path": { "docCount": "_count", "totalDocs": "../total_docs" }, "script": "(docCount / totalDocs) * 100" } } } } } }
返回结果的bySource.buckets中会新增source_percentage字段,直接得到占比数值(比如instagram对应约89%,twitter对应约11%)。
方式二:客户端自行计算占比
如果已经提前获取了总命中文档数(比如你已知的182),可以直接用各来源的doc_count除以总数量再乘以100,手动计算占比:
- Instagram占比:
(162 / 182) * 100 ≈ 89% - Twitter占比:
(20 / 182) * 100 ≈ 11%
这种方式无需修改聚合代码,适合总数量明确且无需ES返回占比的场景。
内容的提问来源于stack exchange,提问作者migueptq
相关产品推荐
相关产品推荐

