You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Elasticsearch聚合中获取各Source的文档占比?

实现Instagram/Twitter文档占比的方案

完全可以实现需求,有两种常用方式:

方式一:在Elasticsearch聚合中直接计算占比

利用Elasticsearch的管道聚合(bucket_script),结合总文档数来计算每个来源的占比。修改后的聚合代码如下:

{
  "aggs": {
    "total_docs": {
      "value_count": {
        "field": "_id"
      }
    },
    "bySource": {
      "terms": {
        "field": "profile.source.keyword",
        "include": ["instagram", "twitter"] // 只统计指定来源,减少不必要计算
      },
      "aggs": {
        "source_percentage": {
          "bucket_script": {
            "buckets_path": {
              "docCount": "_count",
              "totalDocs": "../total_docs"
            },
            "script": "(docCount / totalDocs) * 100"
          }
        }
      }
    }
  }
}

返回结果的bySource.buckets中会新增source_percentage字段,直接得到占比数值(比如instagram对应约89%,twitter对应约11%)。

方式二:客户端自行计算占比

如果已经提前获取了总命中文档数(比如你已知的182),可以直接用各来源的doc_count除以总数量再乘以100,手动计算占比:

  • Instagram占比:(162 / 182) * 100 ≈ 89%
  • Twitter占比:(20 / 182) * 100 ≈ 11%

这种方式无需修改聚合代码,适合总数量明确且无需ES返回占比的场景。

内容的提问来源于stack exchange,提问作者migueptq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 16:55:47