You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Elasticsearch中按嵌套类型排他性键值属性分别统计文档数

问题描述

需要统计Elasticsearch中嵌套类型属性字段里,给定一组排他性name-value属性对各自匹配的文档数量。要求每个属性过滤器单独计数——即使文档满足多个过滤器,每个匹配的过滤器都要计入该属性的文档数;同时要输出所有指定属性的计数,包括匹配数为0的情况。

示例Elasticsearch数据:

{"took":1,"timed_out":false,"_shards":{"total":1,"successful":1,"skipped":0,"failed":0},"hits":{"total":{"value":9,"relation":"eq"},"max_score":1.0020497,"hits":[{"_index":"INDEX","_type":"_doc","_id":"DOC_1","_score":1.0020497,"_source":{"doc":{"attributes":[{"name":"NAME_1","value":"VALUE_1"},{"name":"NAME_2","value":"VALUE_2"},{"name":"NAME_3","value":"VALUE_3"}],"unique_doc_id":"DOC_1"}}},{"_index":"INDEX","_type":"_doc","_id":"DOC_2","_score":1.0020497,"_source":{"doc":{"attributes":[{"name":"NAME_1","value":"VALUE_1"},{"name":"NAME_2","value":"VALUE_7"},{"name":"NAME_4","value":"VALUE_6"}],"unique_doc_id":"DOC_2"}}}]}]

示例输入属性对:

{"attributes":[{"name":"NAME_1","value":"VALUE_1"},{"name":"NAME_2","value":"VALUE_7"},{"name":"NAME_3","value":"VALUE_30"}]}

预期输出:

[{"attribute_name":"NAME_1","count":2},{"attribute_name":"NAME_2","count":1},{"attribute_name":"NAME_3","count":0}]

之前尝试用Painless脚本实现,但核心问题是无法在脚本中维护全局变量来累计跨文档的计数。


解决方案

Elasticsearch的聚合功能更适合这类统计需求,以下提供三种可行方案:

方案一:多搜索(Multi Search)

针对每个属性对单独发起查询,通过_msearch接口批量执行,最后汇总结果。这种方式简单直观,适合属性对数量较少的场景。

构造批量查询

每个查询针对单个name-value属性对,使用nested查询匹配嵌套字段:

{"index": "INDEX"}
{
  "query": {
    "nested": {
      "path": "doc.attributes",
      "query": {
        "bool": {
          "must": [
            {"term": {"doc.attributes.name": "NAME_1"}},
            {"term": {"doc.attributes.value": "VALUE_1"}}
          ]
        }
      }
    }
  },
  "size": 0
}
{"index": "INDEX"}
{
  "query": {
    "nested": {
      "path": "doc.attributes",
      "query": {
        "bool": {
          "must": [
            {"term": {"doc.attributes.name": "NAME_2"}},
            {"term": {"doc.attributes.value": "VALUE_7"}}
          ]
        }
      }
    }
  },
  "size": 0
}
{"index": "INDEX"}
{
  "query": {
    "nested": {
      "path": "doc.attributes",
      "query": {
        "bool": {
          "must": [
            {"term": {"doc.attributes.name": "NAME_3"}},
            {"term": {"doc.attributes.value": "VALUE_30"}}
          ]
        }
      }
    }
  },
  "size": 0
}

结果处理

每个响应中的hits.total.value即为对应属性的匹配文档数,手动整理成预期格式即可,注意给无匹配的属性设置计数为0。

方案二:脚本化度量聚合(Scripted Metric Aggregation)

通过scripted_metric聚合在一次请求中完成所有属性的统计,适合属性对数量较多的场景。该聚合支持分片级别的状态初始化、文档处理和结果合并,完美解决全局计数的问题。

聚合查询

{
  "size": 0,
  "aggs": {
    "attribute_counts": {
      "scripted_metric": {
        "init_script": "state.counts = new HashMap(); state.targets = [['NAME_1', 'VALUE_1'], ['NAME_2', 'VALUE_7'], ['NAME_3', 'VALUE_30']];",
        "map_script": """
          def attributes = doc['doc.attributes'];
          def matchedNames = new HashSet();
          // 收集当前文档匹配的目标属性名称(去重,一个文档同属性只计一次)
          for (def attr in attributes) {
            def name = attr.name.value;
            def value = attr.value.value;
            for (def target in state.targets) {
              if (target[0] == name && target[1] == value) {
                matchedNames.add(name);
                break;
              }
            }
          }
          // 更新计数
          for (def name in matchedNames) {
            state.counts.put(name, (state.counts.getOrDefault(name, 0) as int) + 1);
          }
        """,
        "combine_script": "return state.counts;",
        "reduce_script": """
          def finalCounts = new HashMap();
          // 初始化所有目标属性计数为0
          for (def target in states[0].targets) {
            finalCounts.put(target[0], 0);
          }
          // 合并所有分片的计数结果
          for (def state in states) {
            for (def entry in state.counts.entrySet()) {
              finalCounts.put(entry.getKey(), finalCounts.get(entry.getKey()) + entry.getValue());
            }
          }
          // 转换为预期的数组格式
          def result = new ArrayList();
          for (def entry in finalCounts.entrySet()) {
            result.add(['attribute_name': entry.getKey(), 'count': entry.getValue()]);
          }
          return result;
        """
      }
    }
  }
}

结果提取

查询结果中aggregations.attribute_counts.value直接就是预期的输出数组。

方案三:嵌套+过滤器+反向嵌套聚合

利用nested聚合进入嵌套字段,结合filter聚合匹配属性对,再通过reverse_nested回到根文档,用cardinality统计唯一文档数。适合属性对少、需要避免脚本的场景。

聚合查询

{
  "size": 0,
  "aggs": {
    "nested_attrs": {
      "nested": {
        "path": "doc.attributes"
      },
      "aggs": {
        "NAME_1_count": {
          "filter": {
            "bool": {
              "must": [
                {"term": {"doc.attributes.name": "NAME_1"}},
                {"term": {"doc.attributes.value": "VALUE_1"}}
              ]
            }
          },
          "aggs": {
            "root_docs": {
              "reverse_nested": {},
              "aggs": {
                "count": {
                  "cardinality": {
                    "field": "doc.unique_doc_id"
                  }
                }
              }
            }
          }
        },
        "NAME_2_count": {
          "filter": {
            "bool": {
              "must": [
                {"term": {"doc.attributes.name": "NAME_2"}},
                {"term": {"doc.attributes.value": "VALUE_7"}}
              ]
            }
          },
          "aggs": {
            "root_docs": {
              "reverse_nested": {},
              "aggs": {
                "count": {
                  "cardinality": {
                    "field": "doc.unique_doc_id"
                  }
                }
              }
            }
          }
        },
        "NAME_3_count": {
          "filter": {
            "bool": {
              "must": [
                {"term": {"doc.attributes.name": "NAME_3"}},
                {"term": {"doc.attributes.value": "VALUE_30"}}
              ]
            }
          },
          "aggs": {
            "root_docs": {
              "reverse_nested": {},
              "aggs": {
                "count": {
                  "cardinality": {
                    "field": "doc.unique_doc_id"
                  }
                }
              }
            }
          }
        }
      }
    }
  }
}

结果处理

每个过滤器聚合下的root_docs.count.value即为对应属性的匹配文档数,整理成预期格式即可。


内容的提问来源于stack exchange,提问作者Akshay Agarwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 12:42:39