You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenSearch单文档内嵌套数据聚合与过滤实现咨询

单文档嵌套数据聚合+结果过滤解决方案

针对你的需求——在单文档维度对嵌套数据聚合,再基于聚合结果过滤文档,结合OpenSearch不支持Elasticsearch新版本runtime mapping的限制,可通过以下方式实现:

1. 核心思路

  • 先过滤出同时包含shopId 1和2的文档(满足需求中的前置条件)
  • 使用scripted_metric聚合在每个文档内部计算符合条件的均价
  • 在聚合的reduce阶段过滤出均价>130的结果

2. 完整查询请求

GET warehouse/_search
{
  "size": 0,
  "query": {
    "nested": {
      "path": "inventory",
      "query": {
        "bool": {
          "must": [
            {
              "term": { "inventory.shopId": "1" }
            },
            {
              "term": { "inventory.shopId": "2" }
            }
          ]
        }
      }
    }
  },
  "aggs": {
    "qualified_docs": {
      "scripted_metric": {
        "init_script": "state.sum = 0; state.count = 0; state.doc_id = ''; state.name = ''",
        "map_script": """
          state.doc_id = params._id;
          state.name = params._source.profile.name;
          for (item in params._source.inventory) {
            if (item.shopId == '1' || item.shopId == '2') {
              state.sum += item.price;
              state.count += 1;
            }
          }
        """,
        "combine_script": """
          if (state.count > 0) {
            return {
              'doc_id': state.doc_id,
              'warehouse_name': state.name,
              'avg_price': state.sum / state.count,
              'total_price': state.sum,
              'qualified_item_count': state.count
            };
          } else {
            return null;
          }
        """,
        "reduce_script": """
          def results = [];
          for (res in states) {
            if (res != null && res.avg_price > 130) {
              results.add(res);
            }
          }
          return results;
        """
      }
    }
  }
}

3. 关键部分解释

查询条件

使用nested查询+bool must确保文档同时包含shopId 1和2的库存条目,直接过滤掉不满足前置条件的文档。

Scripted Metric聚合

  • init_script:初始化用于计算的变量(总价、计数、文档ID、仓库名称)
  • map_script:遍历当前文档的inventory数组,筛选shopId为1或2的条目,累加总价和计数
  • combine_script:计算当前文档的均价,返回包含文档标识、聚合结果的对象
  • reduce_script:遍历所有文档的聚合结果,过滤出均价>130的条目,最终返回符合要求的结果列表

4. 文档结构优化建议(可选)

如果这类查询是高频操作,可考虑扁平化文档结构:将每个库存条目作为独立文档,冗余存储仓库的profile信息。例如:

PUT warehouse/_doc/1
{
  "profile": { "name": "Place1" },
  "equipment": "guitar",
  "price": 1000.00,
  "shopId": "1"
}
PUT warehouse/_doc/2
{
  "profile": { "name": "Place1" },
  "equipment": "guitar",
  "price": 200.00,
  "shopId": "2"
}

这种结构下,可通过terms聚合按profile.name分组,再嵌套filter和avg聚合实现需求,查询逻辑更简洁,但会增加数据存储量,需根据业务场景权衡。

内容的提问来源于stack exchange,提问作者John Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 18:14:52