You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch 7.14:嵌套对象数组字段排序及分页方案咨询

Elasticsearch 7.14 嵌套字段排序与分页实现方案

场景说明

使用Elasticsearch 7.14版本,为节省存储空间,将多个共享公共字段的文档合并索引,差异字段存入嵌套对象数组compressedFields。示例原始文档如下:

[
    {
        "uid" : "aaaa",
        "table": "tableA",
        "compressedFields": [
            {
                "date": "28/03/2023 09:50:00",
                "classification": "1", 
                "uid": "12344"
            },
            {
                "date": "28/03/2023 10:00:00",
                "classification": "2", 
                "uid": "55555"
            },
            {
                "date": "28/03/2023 10:10:00",
                "classification": "2", 
                "uid": "6666"
            }
        ]
    },
    {
        "uid" : "bbbbb",
        "table": "tableB",
        "compressedFields": [
            {
                "date": "28/03/2023 09:40:00",
                "classification": "1", 
                "uid": "1111"
            },
            {
                "date": "28/03/2023 09:55:00",
                "classification": "2", 
                "uid": "2222"
            },
            {
                "date": "28/03/2023 10:05:00",
                "classification": "2", 
                "uid": "3333"
            }
        ]
    }
]

需求

检索时按嵌套字段(如date或classification)排序,将嵌套数组展开为扁平文档,预期结果如下:

[
    {
        "uid" : "bbbbb",
        "table": "tableB",
        "date": "28/03/2023 09:40:00",
        "classification": "1",
        "nested_uid": "1111"
    },
    {
        "uid" : "aaaa",
        "table": "tableA",
        "date": "28/03/2023 09:50:00",
        "classification": "1",
        "nested_uid": "12344"
    },
    {
        "uid" : "bbbbb",
        "table": "tableB",
        "date": "28/03/2023 09:55:00",
        "classification": "2",
        "nested_uid": "2222"
    },
    {
        "uid" : "aaaa",
        "table": "tableA",
        "date": "28/03/2023 10:00:00",
        "classification": "2",
        "nested_uid": "55555"
    },
    {
        "uid" : "bbbbb",
        "table": "tableB",
        "date": "28/03/2023 10:05:00",
        "classification": "2",
        "nested_uid": "3333"
    },
    {
        "uid" : "aaaa",
        "table": "tableA",
        "date": "28/03/2023 10:10:00",
        "classification": "2",
        "nested_uid": "6666"
    }
]

注:原预期结果存在重复uid字段,此处调整为nested_uid避免字段冲突。

现有问题

尝试使用聚合查询实现,但该方案无法支持分页,聚合语句如下:

GET my_index/_search
{
    "aggs": {
        "aggs1": {
            "nested": {
                "path": "compressedFields"
            },
            "aggs": {
                "agg2": {
                    "terms": {
                        "field": "compressedFields.date",
                        "order": {
                            "_key": "asc"
                        },
                        "size": 10000
                    },
                    "aggs": {
                        "agg3": {
                            "reverse_nested": {},
                            "aggs": {
                                "agg4": {
                                    "top_hits": {}
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}

最优解决方案

方案1:嵌套查询+Inner Hits+脚本字段(无需重索引)

通过nested查询匹配所有嵌套文档,利用inner_hits获取嵌套字段并指定排序规则,配合脚本字段提取父文档公共字段,最后通过from和size实现分页。

示例查询(按date升序,分页第1页,每页3条):

GET my_index/_search
{
    "query": {
        "nested": {
            "path": "compressedFields",
            "query": {
                "match_all": {}
            },
            "inner_hits": {
                "size": 3,
                "from": 0,
                "sort": [
                    {
                        "compressedFields.date": {
                            "order": "asc",
                            "nested": {
                                "path": "compressedFields"
                            }
                        }
                    }
                ],
                "_source": {
                    "includes": ["compressedFields.*"]
                },
                "script_fields": {
                    "parent_uid": {
                        "script": {
                            "source": "doc['uid'].value"
                        }
                    },
                    "parent_table": {
                        "script": {
                            "source": "doc['table'].value"
                        }
                    }
                }
            }
        }
    },
    "_source": false
}

结果处理

返回的inner_hits.compressedFields.hits.hits中,每个文档包含fields(父文档公共字段)和_source(嵌套字段),在应用层将两者合并即可得到预期的扁平文档结构。

方案2:Ingest Pipeline预处理(重索引为扁平文档)

若允许数据重索引,可通过Ingest Pipeline将嵌套数组拆分为独立扁平文档,从根本上简化排序和分页操作:

  1. 创建拆分Pipeline:
PUT _ingest/pipeline/expand_compressed_fields
{
    "processors": [
        {
            "split": {
                "field": "compressedFields",
                "target_field": "temp_field"
            }
        },
        {
            "script": {
                "source": """
                    ctx.date = ctx.temp_field.date;
                    ctx.classification = ctx.temp_field.classification;
                    ctx.nested_uid = ctx.temp_field.uid;
                    ctx.remove('compressedFields');
                    ctx.remove('temp_field');
                """
            }
        }
    ]
}
  1. 重索引数据到新索引:
POST _reindex
{
    "source": {
        "index": "my_index"
    },
    "dest": {
        "index": "my_index_expanded",
        "pipeline": "expand_compressed_fields"
    }
}

之后直接在新索引上执行普通排序、分页查询即可,无需处理嵌套结构,性能更优。

方案对比

  • 方案1:无需修改现有索引,适配无法重索引的场景,但需应用层处理结果合并,查询性能略低于扁平索引。
  • 方案2:数据结构直观,支持所有Elasticsearch原生功能,性能最优,但需重新索引数据,适合可接受数据迁移的场景。

内容的提问来源于stack exchange,提问作者צחי

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 05:44:54