You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Elasticsearch聚合时如何从数组对应索引获取关联记录?

Elasticsearch数组字段配对聚合解决方案

当cityName和cityId变为同索引对应的数组后,普通的terms聚合无法直接关联两者的对应关系(Elasticsearch会将数组字段扁平化处理,丢失索引关联)。需要使用**脚本化指标聚合(scripted_metric)**来实现按城市名称分桶并获取对应ID的需求,具体聚合语句如下:

{
  "size": 0,
  "aggs": {
    "cName": {
      "scripted_metric": {
        "init_script": "state.cityMap = new HashMap()",
        "map_script": """
          def cityNames = doc['cityName.keyword'].values;
          def cityIds = doc['cityId.keyword'].values;
          // 按数组索引配对名称和ID
          for (int i = 0; i < cityNames.length; i++) {
            def name = cityNames[i];
            def id = cityIds[i];
            // 存储每个城市对应的ID集合
            if (!state.cityMap.containsKey(name)) {
              state.cityMap.put(name, new HashSet());
            }
            state.cityMap.get(name).add(id);
          }
        """,
        "combine_script": "return state.cityMap",
        "reduce_script": """
          def finalCityMap = new HashMap();
          // 合并所有分片的结果
          for (shardMap in states) {
            for (entry in shardMap.entrySet()) {
              def cityName = entry.getKey();
              def ids = entry.getValue();
              if (!finalCityMap.containsKey(cityName)) {
                finalCityMap.put(cityName, new HashSet());
              }
              finalCityMap.get(cityName).addAll(ids);
            }
          }
          // 转换为与原聚合格式一致的输出
          def buckets = [];
          for (entry in finalCityMap.entrySet()) {
            buckets.add([
              "key": entry.getKey(),
              "doc_count": entry.getValue().size(),
              "cId": [
                "buckets": entry.getValue().stream().map(id -> ["key": id, "doc_count": 1]).collect(Collectors.toList())
              ]
            ]);
          }
          return ["buckets": buckets];
        """
      }
    }
  }
}

脚本说明

  • init_script:初始化一个哈希表,用于存储城市名称到对应ID集合的映射
  • map_script:遍历单个文档中的cityName和cityId数组,按索引位置配对两者,将ID存入对应城市名称的集合中
  • combine_script:汇总当前分片内的所有城市映射结果
  • reduce_script:合并所有分片的结果,并转换成和原聚合语句一致的输出格式,方便后续业务兼容

优化建议

如果业务中每个城市名称只会对应唯一的cityId,可以将脚本中的HashSet改为直接存储单个值,减少不必要的集合操作,提升性能:
修改map_script中的存储逻辑:

if (!state.cityMap.containsKey(name)) {
  state.cityMap.put(name, id);
}

同时调整reduce_script中的合并和格式转换逻辑,适配单个值的存储结构。

另外,如果数据量较大,更高效的方案是调整数据结构,将每个城市的名称和ID封装为嵌套对象:

{
  "usr" : "plore113",
  "cities": [
    {"name": "New York", "id": "150p7"},
    {"name": "Delhi", "id": "171x9"}
  ]
}

之后使用nested聚合即可更高效地实现需求,避免脚本带来的性能开销。

内容的提问来源于stack exchange,提问作者Raghav Mishra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 09:52:34