Elasticsearch 7.14:嵌套对象数组字段排序及分页方案咨询
Elasticsearch 7.14 嵌套字段排序与分页实现方案
场景说明
使用Elasticsearch 7.14版本,为节省存储空间,将多个共享公共字段的文档合并索引,差异字段存入嵌套对象数组compressedFields。示例原始文档如下:
[ { "uid" : "aaaa", "table": "tableA", "compressedFields": [ { "date": "28/03/2023 09:50:00", "classification": "1", "uid": "12344" }, { "date": "28/03/2023 10:00:00", "classification": "2", "uid": "55555" }, { "date": "28/03/2023 10:10:00", "classification": "2", "uid": "6666" } ] }, { "uid" : "bbbbb", "table": "tableB", "compressedFields": [ { "date": "28/03/2023 09:40:00", "classification": "1", "uid": "1111" }, { "date": "28/03/2023 09:55:00", "classification": "2", "uid": "2222" }, { "date": "28/03/2023 10:05:00", "classification": "2", "uid": "3333" } ] } ]
需求
检索时按嵌套字段(如date或classification)排序,将嵌套数组展开为扁平文档,预期结果如下:
[ { "uid" : "bbbbb", "table": "tableB", "date": "28/03/2023 09:40:00", "classification": "1", "nested_uid": "1111" }, { "uid" : "aaaa", "table": "tableA", "date": "28/03/2023 09:50:00", "classification": "1", "nested_uid": "12344" }, { "uid" : "bbbbb", "table": "tableB", "date": "28/03/2023 09:55:00", "classification": "2", "nested_uid": "2222" }, { "uid" : "aaaa", "table": "tableA", "date": "28/03/2023 10:00:00", "classification": "2", "nested_uid": "55555" }, { "uid" : "bbbbb", "table": "tableB", "date": "28/03/2023 10:05:00", "classification": "2", "nested_uid": "3333" }, { "uid" : "aaaa", "table": "tableA", "date": "28/03/2023 10:10:00", "classification": "2", "nested_uid": "6666" } ]
注:原预期结果存在重复uid字段,此处调整为nested_uid避免字段冲突。
现有问题
尝试使用聚合查询实现,但该方案无法支持分页,聚合语句如下:
GET my_index/_search { "aggs": { "aggs1": { "nested": { "path": "compressedFields" }, "aggs": { "agg2": { "terms": { "field": "compressedFields.date", "order": { "_key": "asc" }, "size": 10000 }, "aggs": { "agg3": { "reverse_nested": {}, "aggs": { "agg4": { "top_hits": {} } } } } } } } } }
最优解决方案
方案1:嵌套查询+Inner Hits+脚本字段(无需重索引)
通过nested查询匹配所有嵌套文档,利用inner_hits获取嵌套字段并指定排序规则,配合脚本字段提取父文档公共字段,最后通过from和size实现分页。
示例查询(按date升序,分页第1页,每页3条):
GET my_index/_search { "query": { "nested": { "path": "compressedFields", "query": { "match_all": {} }, "inner_hits": { "size": 3, "from": 0, "sort": [ { "compressedFields.date": { "order": "asc", "nested": { "path": "compressedFields" } } } ], "_source": { "includes": ["compressedFields.*"] }, "script_fields": { "parent_uid": { "script": { "source": "doc['uid'].value" } }, "parent_table": { "script": { "source": "doc['table'].value" } } } } } }, "_source": false }
结果处理
返回的inner_hits.compressedFields.hits.hits中,每个文档包含fields(父文档公共字段)和_source(嵌套字段),在应用层将两者合并即可得到预期的扁平文档结构。
方案2:Ingest Pipeline预处理(重索引为扁平文档)
若允许数据重索引,可通过Ingest Pipeline将嵌套数组拆分为独立扁平文档,从根本上简化排序和分页操作:
- 创建拆分Pipeline:
PUT _ingest/pipeline/expand_compressed_fields { "processors": [ { "split": { "field": "compressedFields", "target_field": "temp_field" } }, { "script": { "source": """ ctx.date = ctx.temp_field.date; ctx.classification = ctx.temp_field.classification; ctx.nested_uid = ctx.temp_field.uid; ctx.remove('compressedFields'); ctx.remove('temp_field'); """ } } ] }
- 重索引数据到新索引:
POST _reindex { "source": { "index": "my_index" }, "dest": { "index": "my_index_expanded", "pipeline": "expand_compressed_fields" } }
之后直接在新索引上执行普通排序、分页查询即可,无需处理嵌套结构,性能更优。
方案对比
- 方案1:无需修改现有索引,适配无法重索引的场景,但需应用层处理结果合并,查询性能略低于扁平索引。
- 方案2:数据结构直观,支持所有Elasticsearch原生功能,性能最优,但需重新索引数据,适合可接受数据迁移的场景。
内容的提问来源于stack exchange,提问作者צחי
相关产品推荐
相关产品推荐

