如何使用Elasticsearch筛选器获取每个transaction对应的最新文档并实现年份直方图聚合
Got it, let's tackle this problem step by step. You need two key outcomes: fetching the most recent document for each transaction, then aggregating those results into a yearly histogram. A basic terms aggregation won't cut it for the yearly breakdown (since it requires manual year listing), so here's a scalable, one-query solution using Elasticsearch aggregations:
1. Get the Latest Document per Transaction
First, we'll group documents by the transaction field and pull only the most recent entry in each group using a combination of terms and top_hits aggregations:
{ "size": 0, "aggs": { "group_by_transaction": { "terms": { "field": "transaction.keyword" }, "aggs": { "latest_document": { "top_hits": { "size": 1, "sort": [ { "date": { "order": "desc" } } ] } }, "latest_date": { "max": { "field": "date" } } } } } }
size: 0: Skips returning raw hits since we focus on aggregation results.group_by_transaction: Groups all documents by exacttransactionvalues (using.keywordto avoid text analysis issues).latest_document: Usestop_hitsto fetch only the single most recent document in each group, sorted bydatedescending. This gives you your target results: transaction C (1999-10-26), A (1999-12-25), B (2000-10-25).latest_date: Captures the most recent date for each transaction (we'll use this for the yearly aggregation next).
2. Yearly Histogram Aggregation
To get the yearly count without manual year entries, we'll add a pipeline date_histogram aggregation that operates on the latest_date values from the first step. This automatically bins the dates into yearly intervals and counts how many transactions fall into each year:
{ "size": 0, "aggs": { "group_by_transaction": { "terms": { "field": "transaction.keyword" }, "aggs": { "latest_document": { "top_hits": { "size": 1, "sort": [{"date": "desc"}] } }, "latest_date": { "max": { "field": "date" } } } }, "yearly_transaction_count": { "date_histogram": { "field": "_value", "interval": "year", "format": "yyyy" }, "buckets_path": "group_by_transaction>latest_date" } } }
yearly_transaction_count: This pipeline aggregation takes thelatest_datevalues from each transaction bucket, bins them into yearly intervals (interval: "year"), and formats the bucket keys as 4-digit years (format: "yyyy").- The result will be exactly your expected histogram:
{ "1999": 2, "2000": 1 }(since transactions A and C have latest dates in 1999, B in 2000).
Key Advantages
- No manual year listing: The
date_histogramautomatically handles any date range, making it scalable for large datasets. - Single query: You get both the latest documents per transaction and the yearly count in one request.
内容的提问来源于stack exchange,提问作者BenjaminC

