You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从ElasticSearch每个唯一Tagname桶中获取Top5文档

How to Get Top 5 Documents per Unique Tagname Bucket in Elasticsearch

Got it, let's break down how to pull off this requirement—getting the top 5 documents for each unique Tagname (which is mapped as a keyword type). The key here is combining two Elasticsearch aggregations: terms (to group by tag) and top_hits (to fetch the top docs in each group).

Step-by-Step Solution

Here's the complete DSL query you can use:

{
  "size": 0, // We don't need top-level hits, just the aggregations
  "aggs": {
    "tag_buckets": {
      "terms": {
        "field": "Tagname", // Group by the keyword field Tagname
        "size": 10000 // Adjust this if you have more than 10k unique tags
      },
      "aggs": {
        "top_5_docs": {
          "top_hits": {
            "size": 5, // Return top 5 docs per tag bucket
            "_source": ["Tagname", "Title"], // Specify fields to include (optional)
            "sort": [ // Optional: Add sorting if needed, e.g., by relevance or a custom field
              { "_score": { "order": "desc" } }
            ]
          }
        }
      }
    }
  }
}

Let's explain each part:

  • size: 0: We don't care about the top-level search hits since we're only interested in the aggregated results. This saves bandwidth and processing time.
  • terms aggregation (tag_buckets): Groups all documents by the Tagname field. The size parameter here controls how many unique tags we want to return—set it higher if you have more than 10k unique tags.
  • top_hits aggregation (top_5_docs): Nested inside the terms aggregation, this fetches the top 5 documents for each tag bucket. You can customize:
    • size: Explicitly set to 5 to get the top 5 docs.
    • _source: Pick specific fields to return instead of the full document (great for reducing payload size).
    • sort: Add sorting rules if you want to prioritize docs by something other than relevance score—for example, if you had a timestamp field, you could sort by that to get the newest 5 docs per tag.

Example Response Structure

The response will look something like this, with each tag bucket containing its top 5 documents:

{
  "aggregations": {
    "tag_buckets": {
      "buckets": [
        {
          "key": "Veniam",
          "doc_count": 42,
          "top_5_docs": {
            "hits": {
              "total": { "value": 42, "relation": "eq" },
              "hits": [
                {
                  "_source": { "Tagname": ["Veniam"], "Title": ["Occaecat do. Eu ut."] }
                },
                // 4 more documents for "Veniam"
              ]
            }
          }
        },
        // More tag buckets with their top 5 docs
      ]
    }
  }
}

Notes to Keep in Mind

  • If you need to sort the documents in each bucket by a specific field (not just relevance), update the sort array in the top_hits aggregation. For example, to sort by a hypothetical created_at field in descending order:
    "sort": [ { "created_at": { "order": "desc" } } ]
    
  • The terms aggregation's size parameter defaults to 10, so make sure to set it to a value that covers all your unique tags if you need to include every single one.

内容的提问来源于stack exchange,提问作者Temp O'rary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:41:44