Elasticsearch Transform中嵌套字段聚合的正确实现咨询
场景说明
我用elasticsearch-dsl定义了包含嵌套字段products的ProductCategory文档,代码如下:
class ProductCategory(Document): category_id = Integer(required=True) products = Nested(Products, multi=True) date = Date(required=True) class Products(InnerDoc): """ Class used to represent a denormalized user stored on other objects. """ product_id = Integer(multi=True) category_position = Integer(multi=True)
尝试在Elasticsearch Transform中对该嵌套字段执行聚合时,初始配置运行报错,临时添加外层聚合后看似可行,想确认该方法是否正确,以及嵌套字段聚合的正确方式。
初始Transform配置及错误
初始的Transform配置如下:
{ "source": { "index": ["category_index"], "query": { "bool": { "filter": [ { "range": { "date": { "gte": "now-365d", "lte": "now" } } } ] } } }, "dest": { "index": "category_transformation" }, "sync": { "time": { "field": "date", "delay": "1h" } }, "pivot": { "group_by": { "category_id": { "terms": { "field": "category_id" } } }, "aggs": { "related_product_info": { "nested": { "path": "products" }, "aggs": { "product_info": { "filter": { "range": { "products.category_position": { "gte": 1, "lte": 6000 } } } } } } } } }
运行时触发错误:
Validation Failed: 1: Unsupported aggregation type [nested]
临时解决方案配置
为了绕开错误,我在嵌套聚合外层添加了一层terms聚合,配置如下:
{ "source": { "index": ["category_index"], "query": { "bool": { "filter": [ { "range": { "date": { "gte": "now-365d", "lte": "now" } } } ] } } }, "dest": { "index": "category_transformation" }, "sync": { "time": { "field": "date", "delay": "1h" } }, "pivot": { "group_by": { "category_id": { "terms": { "field": "category_id" } } }, "aggregations": { "nested_category_info": { "terms": { "field": "category_id", "size": 1 }, "aggs": { "related_product_info": { "nested": { "path": "products" }, "aggs": { "product_info": { "filter": { "range": { "products.category_position": { "gte": 1, "lte": 6000 } } } } } } } } } } }
该配置看似能运行,但不确定是否正确。
问题
请问这个临时方案是否正确?如果不正确,Elasticsearch Transform中对嵌套字段执行聚合的正确方式是什么?
解答
临时方案的问题
这个临时方案不正确:外层额外添加的terms聚合(按category_id分组,size=1)完全冗余,因为pivot层已经按category_id完成了分组,重复分组会导致结果结构冗余,还会降低聚合效率。
正确实现方式
Elasticsearch Transform的pivot不允许将nested聚合直接作为顶级聚合,但可以通过以下两种规范方式实现嵌套字段聚合:
方式1:直接嵌套nested聚合(适配较新版本Elasticsearch)
如果你的Elasticsearch版本支持将nested作为子聚合直接放在aggregations下,可以直接配置:
{ "source": { "index": ["category_index"], "query": { "bool": { "filter": [ { "range": { "date": { "gte": "now-365d", "lte": "now" } } } ] } } }, "dest": { "index": "category_transformation" }, "sync": { "time": { "field": "date", "delay": "1h" } }, "pivot": { "group_by": { "category_id": { "terms": { "field": "category_id" } } }, "aggregations": { "filtered_products": { "nested": { "path": "products" }, "aggs": { "position_filter": { "filter": { "range": { "products.category_position": { "gte": 1, "lte": 6000 } } }, "aggs": { // 按需添加具体聚合逻辑,比如统计去重商品ID数量 "unique_product_count": { "cardinality": { "field": "products.product_id" } } } } } } } } }
方式2:用match_all过滤聚合包裹(适配旧版本Elasticsearch)
如果旧版本不允许直接将nested作为子聚合,可使用match_all的过滤聚合作为外层,内部嵌套nested聚合:
{ "source": { "index": ["category_index"], "query": { "bool": { "filter": [ { "range": { "date": { "gte": "now-365d", "lte": "now" } } } ] } } }, "dest": { "index": "category_transformation" }, "sync": { "time": { "field": "date", "delay": "1h" } }, "pivot": { "group_by": { "category_id": { "terms": { "field": "category_id" } } }, "aggregations": { "root_wrapper": { "filter": { "match_all": {} }, "aggs": { "related_product_info": { "nested": { "path": "products" }, "aggs": { "product_info": { "filter": { "range": { "products.category_position": { "gte": 1, "lte": 6000 } } }, "aggs": { "top_products": { "top_hits": { "size": 10, "_source": ["products.product_id", "products.category_position"] } } } } } } } } } } }
以上两种方式都能正确实现嵌套字段的聚合逻辑,且没有冗余分组操作,结果结构更清晰。
内容的提问来源于stack exchange,提问作者Lars Hurkmans

