You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Elasticsearch首步聚合结果执行二次聚合统计?

合并两步Elasticsearch聚合逻辑的解决方案

问题说明

先看单条数据结构:

{
  "id": "id123",
  "sc": 1,
  "src": {
    "aid": "aid123",
    "result": {      
      "cid": "cid123",
      "scm": "GL",
      "title": "DTitle",
      "reason": "DReason",
      "status": "Failed",
      "entity": "DEnt"
    },
    "createdAt": 1718864490635,
    "rel": {
      "nam": "dRel",
      "par": "pid123"
    }
  }
}

需求分两步:

  1. 按cid+aid组合分组,获取每组内createdAt最新的条目
  2. 基于第一步的结果,再按cid聚合,统计:
    • 该cid下的总条目数(即不同aid的数量)
    • 状态为Passed的条目数
    • 状态为Failed的条目数
    • 保留result内的所有字段样本

你已经有两个单独运行正常的查询,现在需要把它们合并成一个查询。

合并后的完整查询

GET /index/_search
{
  "size": 0,
  "aggs": {
    "by_complianceId": {
      "terms": {
        "field": "result.cid.keyword"
      },
      "aggs": {
        // 第一步:按aid分组,获取每个cid+aid的最新条目
        "artifact_group": {
          "terms": {
            "field": "aid.keyword"
          },
          "aggs": {
            "latest_execution": {
              "top_hits": {
                "sort": [
                  {
                    "createdAt": {
                      "order": "desc"
                    }
                  }
                ],
                "_source": {
                  "includes": [
                    "result.cid", 
                    "aid", 
                    "result.scm",
                    "result.title",
                    "result.reason",
                    "result.entity",
                    "result.status"
                  ]
                },
                "size": 1
              }
            },
            // 标记当前aid的最新条目是否为Passed
            "is_passed": {
              "filter": {
                "term": {
                  "result.status.keyword": "PASSED"
                }
              }
            },
            // 标记当前aid的最新条目是否为Failed
            "is_failed": {
              "filter": {
                "term": {
                  "result.status.keyword": "FAILED"
                }
              }
            }
          }
        },
        // 统计总条目数:即当前cid下不同aid的数量
        "total_count": {
          "cardinality": {
            "field": "aid.keyword"
          }
        },
        // 统计Passed状态的条目数:求和所有aid组中标记为Passed的数量
        "passed_count": {
          "sum_bucket": {
            "buckets_path": "artifact_group>is_passed._count"
          }
        },
        // 统计Failed状态的条目数:求和所有aid组中标记为Failed的数量
        "failed_count": {
          "sum_bucket": {
            "buckets_path": "artifact_group>is_failed._count"
          }
        },
        // 保留一份result字段的样本文档(取当前cid下最新的一条)
        "sample_result": {
          "top_hits": {
            "sort": [{"createdAt": "desc"}],
            "_source": {
              "includes": [
                "result.cid", 
                "aid", 
                "result.scm",
                "result.title",
                "result.reason",
                "result.entity",
                "result.status"
              ]
            },
            "size": 1
          }
        }
      }
    }
  }
}

逻辑说明

  1. 第一层聚合:按result.cid.keyword分组,把相同合规ID的文档归为一组
  2. 第二层聚合:
    • 按aid.keyword分组,确保每个artifact ID单独成组
    • 用top_hits取每组内createdAt最新的条目,同时指定返回需要的字段
    • 用filter聚合标记该最新条目的状态是否为Passed/Failed,方便后续统计
  3. 顶层统计:
    • total_count用cardinality统计当前cid下不同aid的数量,即第一步得到的总条目数
    • passed_count和failed_count用sum_bucket聚合,把子桶中标记为对应状态的计数求和,得到最终的状态统计数
    • sample_result用top_hits保留当前cid下的一份样本文档,包含所有需要的result字段

注意事项

  • 所有用于分组的字段都加了.keyword后缀,避免分词导致分组错误(比如cid如果是文本类型,分词后会拆分成多个词,导致分组混乱)
  • sum_bucket聚合要求Elasticsearch版本在6.4及以上,如果你的版本低于这个,需要升级或者改用scripted_metric聚合实现统计
  • 如果需要保留所有cid+aid的最新result字段,可以查看artifact_group下的latest_execution结果,里面包含每个aid组的最新条目

内容的提问来源于stack exchange,提问作者Kartik Saurya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 17:49:56