You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch聚合分页:from/size超限及Scroll API问题求助

解决AWS Elasticsearch 6.2聚合分页与Elasticsearch-DSL的问题

从Solr转到AWS托管ES 6.2后,你遇到的聚合分页限制、Scroll API的局限性以及Elasticsearch-DSL的使用问题,都是迁移过程中常见的痛点,我来逐个帮你拆解解决:


问题1:聚合查询中top_hits能否使用from=100?如何实现?

首先明确:top_hits聚合内的from/size是针对每个聚合桶内的结果分页,和全局搜索的from限制(10000)无关。你之前的报错大概率是JSON语法不规范(from/size未加引号)或者对参数的误解,以下是正确的实现方案:

原生Elasticsearch查询修正

修正JSON语法并确保参数正确的查询如下,可直接在Kibana中运行:

GET /properties/_search
{
  "size": 10,
  "query": {
    "bool": {
      "must": [
        { "match": { "published": true } },
        { "match": { "country": "South Africa" } }
      ]
    }
  },
  "aggs": {
    "aggs_by_feed": {
      "terms": {
        "field": "feed",
        "order": { "_key": "desc" }
      },
      "aggs": {
        "tops": {
          "top_hits": {
            "from": 100,
            "size": 50,
            "_source": [ "id", "feed_provider_id" ],
            "sort": [ { "instant_book": { "order": "desc" } } ]
          }
        }
      }
    }
  },
  "sort": [ { "instant_book": { "order": "desc" } } ]
}

Elasticsearch-DSL实现top_hits指定_source

你之前的DSL代码错误在于误用了field参数,top_hits需要通过_source指定返回字段,同时用from_(避开Python关键字)设置分页起始位置:

from elasticsearch_dsl import Search, A

es = init_es()
s = Search(using=es, index=index_name, doc_type=doc_type)

# 构建查询条件
s = s.query("bool", must=[
    {"match": {"published": True}},
    {"match": {"country": "South Africa"}}
])

# 构建聚合:按feed分组,每个组取第101-150条top数据
aggs_by_feed = A('terms', field='feed', order={"_key": "desc"})
top_hits_agg = A(
    'top_hits', 
    from_=100, 
    size=50, 
    _source=["id", "feed_provider_id"], 
    sort=[{"instant_book": {"order": "desc"}}]
)
aggs_by_feed.metric('tops', top_hits_agg)
s.aggs.bucket('aggs_by_feed', aggs_by_feed)

# 执行查询并处理结果
response = s.execute()
print('Hit........')
for hit in response:
    print(hit.meta.score, hit.feed)

print('AGG........')
for bucket in response.aggregations.aggs_by_feed.buckets:
    print(f"Feed: {bucket.key}")
    for top_hit in bucket.tops.hits.hits:
        print(f"ID: {top_hit['_source']['id']}, Feed Provider ID: {top_hit['_source']['feed_provider_id']}")

问题2:如果top_hits分页受限,替代Solr分组功能的方案

若AWS ES 6.2确实限制了top_hits的from值,可通过以下两种方案实现类似Solr的分组分页:

方案1:Composite聚合(推荐,ES 6+原生支持)

Composite聚合支持对聚合桶进行分页,适合处理大量分组的场景,避免一次性返回所有桶:

GET /properties/_search
{
  "size": 0,
  "query": {
    "bool": {
      "must": [
        {"match": {"published": true}},
        {"match": {"country": "South Africa"}}
      ]
    }
  },
  "aggs": {
    "composite_feed": {
      "composite": {
        "size": 10, // 每页返回10个feed分组
        "sources": [
          {"feed": {"terms": {"field": "feed", "order": "desc"}}}
        ]
      },
      "aggs": {
        "tops": {
          "top_hits": {
            "size": 50,
            "_source": ["id", "feed_provider_id"],
            "sort": [{"instant_book": {"order": "desc"}}]
          }
        }
      }
    }
  }
}

获取下一页时,只需在查询中添加after参数,值为上一次返回的composite_feed.after_key:

GET /properties/_search
{
  "size": 0,
  "query": {...}, // 复用之前的查询条件
  "aggs": {
    "composite_feed": {
      "composite": {
        "size": 10,
        "sources": [{"feed": {"terms": {"field": "feed", "order": "desc"}}}],
        "after": {"feed": "上一页最后一个feed值"}
      },
      "aggs": {...} // 复用之前的top_hits聚合
    }
  }
}

方案2:Terms聚合的Partition参数

将所有feed分组拆分为多个分区,每次查询一个分区,适合分组数量极大的场景:

GET /properties/_search
{
  "size": 0,
  "query": {
    "bool": {
      "must": [
        {"match": {"published": true}},
        {"match": {"country": "South Africa"}}
      ]
    }
  },
  "aggs": {
    "aggs_by_feed": {
      "terms": {
        "field": "feed",
        "order": {"_key": "desc"},
        "include": {"partition": 0, "num_partitions": 10} // 分成10个分区,取第0个
      },
      "aggs": {
        "tops": {
          "top_hits": {
            "size": 50,
            "_source": ["id", "feed_provider_id"],
            "sort": [{"instant_book": {"order": "desc"}}]
          }
        }
      }
    }
  }
}

循环修改partition值(从0到9)即可遍历所有feed分组的top数据。


关于Scroll API的补充说明

你提到Scroll API后续调用仅返回Hits是正常现象。Scroll API的设计目的是高效遍历所有文档,聚合结果只会在初始查询中计算一次,后续scroll请求仅返回文档命中,不会重复计算聚合。如果需要分页获取聚合结果,应使用上述Composite或Partition方案,而非Scroll。


内容的提问来源于stack exchange,提问作者A l w a y s S u n n y

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:09:05