从ElasticSearch每个唯一Tagname桶中获取Top5文档
How to Get Top 5 Documents per Unique Tagname Bucket in Elasticsearch
Got it, let's break down how to pull off this requirement—getting the top 5 documents for each unique Tagname (which is mapped as a keyword type). The key here is combining two Elasticsearch aggregations: terms (to group by tag) and top_hits (to fetch the top docs in each group).
Step-by-Step Solution
Here's the complete DSL query you can use:
{ "size": 0, // We don't need top-level hits, just the aggregations "aggs": { "tag_buckets": { "terms": { "field": "Tagname", // Group by the keyword field Tagname "size": 10000 // Adjust this if you have more than 10k unique tags }, "aggs": { "top_5_docs": { "top_hits": { "size": 5, // Return top 5 docs per tag bucket "_source": ["Tagname", "Title"], // Specify fields to include (optional) "sort": [ // Optional: Add sorting if needed, e.g., by relevance or a custom field { "_score": { "order": "desc" } } ] } } } } } }
Let's explain each part:
size: 0: We don't care about the top-level search hits since we're only interested in the aggregated results. This saves bandwidth and processing time.termsaggregation (tag_buckets): Groups all documents by theTagnamefield. Thesizeparameter here controls how many unique tags we want to return—set it higher if you have more than 10k unique tags.top_hitsaggregation (top_5_docs): Nested inside thetermsaggregation, this fetches the top 5 documents for each tag bucket. You can customize:size: Explicitly set to 5 to get the top 5 docs._source: Pick specific fields to return instead of the full document (great for reducing payload size).sort: Add sorting rules if you want to prioritize docs by something other than relevance score—for example, if you had atimestampfield, you could sort by that to get the newest 5 docs per tag.
Example Response Structure
The response will look something like this, with each tag bucket containing its top 5 documents:
{ "aggregations": { "tag_buckets": { "buckets": [ { "key": "Veniam", "doc_count": 42, "top_5_docs": { "hits": { "total": { "value": 42, "relation": "eq" }, "hits": [ { "_source": { "Tagname": ["Veniam"], "Title": ["Occaecat do. Eu ut."] } }, // 4 more documents for "Veniam" ] } } }, // More tag buckets with their top 5 docs ] } } }
Notes to Keep in Mind
- If you need to sort the documents in each bucket by a specific field (not just relevance), update the
sortarray in thetop_hitsaggregation. For example, to sort by a hypotheticalcreated_atfield in descending order:"sort": [ { "created_at": { "order": "desc" } } ] - The
termsaggregation'ssizeparameter defaults to 10, so make sure to set it to a value that covers all your unique tags if you need to include every single one.
内容的提问来源于stack exchange,提问作者Temp O'rary
相关产品推荐
相关产品推荐

