Elasticsearch中如何滚动inner_hits结果无需调整max_inner_result_window?
Elasticsearch inner_hits 全量遍历方案说明
首先明确:Elasticsearch 目前没有提供类似 scroll 或 search_after 的原生机制,专门用来分页遍历单个文档下匹配的嵌套 inner_hits 结果以绕过 max_inner_result_window 限制。以下是几种可行的替代方案:
方案1:临时调整 max_inner_result_window
虽然你希望避免修改配置,但官方设计该参数的目的就是控制 inner_hits 的返回上限。如果是临时业务需求,可以针对性调整索引配置,用完后再恢复默认值:
PUT /your_index/_settings { "index.max_inner_result_window": 10000 }
该方案操作简单,适合偶尔需要全量获取 inner_hits 的场景。
方案2:将嵌套结构转为父子文档(Join 类型)
如果嵌套数据量极大且需要频繁全量遍历匹配项,建议把 nested 类型改为 Join 类型的父子文档结构。这样查询子文档时,就能用普通的 scroll 或 search_after 全量遍历,完全不受 inner_hits 的窗口限制。
示例 mapping 定义:
PUT /your_index { "mappings": { "properties": { "join_field": { "type": "join", "relations": { "parent_doc": "nested_child" } } } } }
查询子文档并使用 scroll 遍历:
POST /your_index/_search?scroll=1m { "query": { "has_parent": { "parent_type": "parent_doc", "query": { // 原父文档的过滤条件 "match": { "parent_field": "target_value" } } } }, "size": 1000 }
方案3:客户端分批过滤(仅适合小数据量场景)
如果匹配的 inner_hits 总数未超出合理范围,只是超过默认窗口,可以先将 inner_hits 的 size 设为当前 max_inner_result_window 的最大值,拿到一批结果后,在客户端记录已获取嵌套项的唯一标识(如每个嵌套项的 id 字段),后续查询时通过 must_not 过滤已返回的项,循环直到无新结果。
第一次查询示例:
POST /your_index/_search { "query": { "nested": { "path": "your_nested_field", "query": { "match": { "your_nested_field.target_field": "value" } }, "inner_hits": { "size": 1000 } } } }
后续查询示例(过滤已获取的项):
POST /your_index/_search { "query": { "nested": { "path": "your_nested_field", "query": { "bool": { "must": [ { "match": { "your_nested_field.target_field": "value" } } ], "must_not": [ { "terms": { "your_nested_field.id": ["已获取的ID列表"] } } ] } }, "inner_hits": { "size": 1000 } } } }
该方案缺点是多次查询会增加开销,且依赖嵌套项有唯一标识。
内容的提问来源于stack exchange,提问作者Tom
相关产品推荐
相关产品推荐

