在Elasticsearch聚合时如何从数组对应索引获取关联记录?
Elasticsearch数组字段配对聚合解决方案
当cityName和cityId变为同索引对应的数组后,普通的terms聚合无法直接关联两者的对应关系(Elasticsearch会将数组字段扁平化处理,丢失索引关联)。需要使用**脚本化指标聚合(scripted_metric)**来实现按城市名称分桶并获取对应ID的需求,具体聚合语句如下:
{ "size": 0, "aggs": { "cName": { "scripted_metric": { "init_script": "state.cityMap = new HashMap()", "map_script": """ def cityNames = doc['cityName.keyword'].values; def cityIds = doc['cityId.keyword'].values; // 按数组索引配对名称和ID for (int i = 0; i < cityNames.length; i++) { def name = cityNames[i]; def id = cityIds[i]; // 存储每个城市对应的ID集合 if (!state.cityMap.containsKey(name)) { state.cityMap.put(name, new HashSet()); } state.cityMap.get(name).add(id); } """, "combine_script": "return state.cityMap", "reduce_script": """ def finalCityMap = new HashMap(); // 合并所有分片的结果 for (shardMap in states) { for (entry in shardMap.entrySet()) { def cityName = entry.getKey(); def ids = entry.getValue(); if (!finalCityMap.containsKey(cityName)) { finalCityMap.put(cityName, new HashSet()); } finalCityMap.get(cityName).addAll(ids); } } // 转换为与原聚合格式一致的输出 def buckets = []; for (entry in finalCityMap.entrySet()) { buckets.add([ "key": entry.getKey(), "doc_count": entry.getValue().size(), "cId": [ "buckets": entry.getValue().stream().map(id -> ["key": id, "doc_count": 1]).collect(Collectors.toList()) ] ]); } return ["buckets": buckets]; """ } } } }
脚本说明
- init_script:初始化一个哈希表,用于存储城市名称到对应ID集合的映射
- map_script:遍历单个文档中的
cityName和cityId数组,按索引位置配对两者,将ID存入对应城市名称的集合中 - combine_script:汇总当前分片内的所有城市映射结果
- reduce_script:合并所有分片的结果,并转换成和原聚合语句一致的输出格式,方便后续业务兼容
优化建议
如果业务中每个城市名称只会对应唯一的cityId,可以将脚本中的HashSet改为直接存储单个值,减少不必要的集合操作,提升性能:
修改map_script中的存储逻辑:
if (!state.cityMap.containsKey(name)) { state.cityMap.put(name, id); }
同时调整reduce_script中的合并和格式转换逻辑,适配单个值的存储结构。
另外,如果数据量较大,更高效的方案是调整数据结构,将每个城市的名称和ID封装为嵌套对象:
{ "usr" : "plore113", "cities": [ {"name": "New York", "id": "150p7"}, {"name": "Delhi", "id": "171x9"} ] }
之后使用nested聚合即可更高效地实现需求,避免脚本带来的性能开销。
内容的提问来源于stack exchange,提问作者Raghav Mishra
相关产品推荐
相关产品推荐

