Elasticsearch聚合分页:from/size超限及Scroll API问题求助
从Solr转到AWS托管ES 6.2后,你遇到的聚合分页限制、Scroll API的局限性以及Elasticsearch-DSL的使用问题,都是迁移过程中常见的痛点,我来逐个帮你拆解解决:
问题1:聚合查询中top_hits能否使用from=100?如何实现?
首先明确:top_hits聚合内的from/size是针对每个聚合桶内的结果分页,和全局搜索的from限制(10000)无关。你之前的报错大概率是JSON语法不规范(from/size未加引号)或者对参数的误解,以下是正确的实现方案:
原生Elasticsearch查询修正
修正JSON语法并确保参数正确的查询如下,可直接在Kibana中运行:
GET /properties/_search { "size": 10, "query": { "bool": { "must": [ { "match": { "published": true } }, { "match": { "country": "South Africa" } } ] } }, "aggs": { "aggs_by_feed": { "terms": { "field": "feed", "order": { "_key": "desc" } }, "aggs": { "tops": { "top_hits": { "from": 100, "size": 50, "_source": [ "id", "feed_provider_id" ], "sort": [ { "instant_book": { "order": "desc" } } ] } } } } }, "sort": [ { "instant_book": { "order": "desc" } } ] }
Elasticsearch-DSL实现top_hits指定_source
你之前的DSL代码错误在于误用了field参数,top_hits需要通过_source指定返回字段,同时用from_(避开Python关键字)设置分页起始位置:
from elasticsearch_dsl import Search, A es = init_es() s = Search(using=es, index=index_name, doc_type=doc_type) # 构建查询条件 s = s.query("bool", must=[ {"match": {"published": True}}, {"match": {"country": "South Africa"}} ]) # 构建聚合:按feed分组,每个组取第101-150条top数据 aggs_by_feed = A('terms', field='feed', order={"_key": "desc"}) top_hits_agg = A( 'top_hits', from_=100, size=50, _source=["id", "feed_provider_id"], sort=[{"instant_book": {"order": "desc"}}] ) aggs_by_feed.metric('tops', top_hits_agg) s.aggs.bucket('aggs_by_feed', aggs_by_feed) # 执行查询并处理结果 response = s.execute() print('Hit........') for hit in response: print(hit.meta.score, hit.feed) print('AGG........') for bucket in response.aggregations.aggs_by_feed.buckets: print(f"Feed: {bucket.key}") for top_hit in bucket.tops.hits.hits: print(f"ID: {top_hit['_source']['id']}, Feed Provider ID: {top_hit['_source']['feed_provider_id']}")
问题2:如果top_hits分页受限,替代Solr分组功能的方案
若AWS ES 6.2确实限制了top_hits的from值,可通过以下两种方案实现类似Solr的分组分页:
方案1:Composite聚合(推荐,ES 6+原生支持)
Composite聚合支持对聚合桶进行分页,适合处理大量分组的场景,避免一次性返回所有桶:
GET /properties/_search { "size": 0, "query": { "bool": { "must": [ {"match": {"published": true}}, {"match": {"country": "South Africa"}} ] } }, "aggs": { "composite_feed": { "composite": { "size": 10, // 每页返回10个feed分组 "sources": [ {"feed": {"terms": {"field": "feed", "order": "desc"}}} ] }, "aggs": { "tops": { "top_hits": { "size": 50, "_source": ["id", "feed_provider_id"], "sort": [{"instant_book": {"order": "desc"}}] } } } } } }
获取下一页时,只需在查询中添加after参数,值为上一次返回的composite_feed.after_key:
GET /properties/_search { "size": 0, "query": {...}, // 复用之前的查询条件 "aggs": { "composite_feed": { "composite": { "size": 10, "sources": [{"feed": {"terms": {"field": "feed", "order": "desc"}}}], "after": {"feed": "上一页最后一个feed值"} }, "aggs": {...} // 复用之前的top_hits聚合 } } }
方案2:Terms聚合的Partition参数
将所有feed分组拆分为多个分区,每次查询一个分区,适合分组数量极大的场景:
GET /properties/_search { "size": 0, "query": { "bool": { "must": [ {"match": {"published": true}}, {"match": {"country": "South Africa"}} ] } }, "aggs": { "aggs_by_feed": { "terms": { "field": "feed", "order": {"_key": "desc"}, "include": {"partition": 0, "num_partitions": 10} // 分成10个分区,取第0个 }, "aggs": { "tops": { "top_hits": { "size": 50, "_source": ["id", "feed_provider_id"], "sort": [{"instant_book": {"order": "desc"}}] } } } } } }
循环修改partition值(从0到9)即可遍历所有feed分组的top数据。
关于Scroll API的补充说明
你提到Scroll API后续调用仅返回Hits是正常现象。Scroll API的设计目的是高效遍历所有文档,聚合结果只会在初始查询中计算一次,后续scroll请求仅返回文档命中,不会重复计算聚合。如果需要分页获取聚合结果,应使用上述Composite或Partition方案,而非Scroll。
内容的提问来源于stack exchange,提问作者A l w a y s S u n n y

