Elasticsearch 7.6.2实现distinct查询及聚合报错解决咨询
Hey there! Let's work through fixing your error and getting exactly the data you need in Elasticsearch 7.6.2.
First, why you're seeing that error
The error message is pretty straightforward once you parse it: your trace_id field is a text type, which Elasticsearch optimizes for full-text search—not operations like aggregations or sorting that need per-document field data. By default, these operations are disabled for text fields to save memory and avoid performance hits.
Fixing the aggregation error (the right way)
The best practice here is to use a keyword field for aggregations. If you didn’t customize your index mapping, Elasticsearch automatically creates a trace_id.keyword sub-field for any text field (it’s a non-analyzed version of the field, perfect for grouping distinct values).
If for some reason that sub-field doesn’t exist, you can add it to your index mapping with this command (note: this won’t retroactively update existing documents—you’ll need to reindex old data if you want it included in aggregations):
PUT /stpjoblogs/_mapping { "properties": { "trace_id": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256 } } } } }
Avoid enabling fielddata on the text field—the error mentions this as an option, but it loads inverted index data into memory which can cause significant performance issues, especially with large datasets. Stick to the keyword sub-field for stability.
Implementing your actual requirement
Your goal is to get the most recent n distinct trace_ids, along with their corresponding job_name and timestamp—mirroring that SQL query. To do this, we’ll combine a terms aggregation (for deduplication) with a top_hits aggregation (to grab the latest record for each trace_id).
Here's the full query (replace 10 with your desired n value):
GET /stpjoblogs/_search { "size": 0, // We don't need raw search results, just aggregation data "query": { "bool": { "must": [{"match": {"status": "SUCCESS"}}] } }, "aggs": { "distinct_trace_ids": { "terms": { "field": "trace_id.keyword", "size": 10, // Number of distinct trace_ids to return "order": { "latest_timestamp": "desc" // Sort buckets by the most recent timestamp } }, "aggs": { "latest_timestamp": { "max": { "field": "timestamp" // Track the newest timestamp per trace_id } }, "latest_record": { "top_hits": { "size": 1, "sort": [{"timestamp": {"order": "desc"}}], "_source": ["trace_id", "job_name", "timestamp"] // Fields to include in results } } } } } }
Breaking down the query:
size: 0: Skips returning raw search documents since we only care about the aggregation output.termsaggregation: Groups results bytrace_id.keywordto get distinct values, and limits the output tonentries.maxaggregation: Calculates the newest timestamp for eachtrace_idbucket, so we can sort the buckets from most recent to oldest.top_hitsaggregation: Pulls the single most recent record (sorted by timestamp) for eachtrace_id, including only the fields you need (trace_id,job_name,timestamp).
What the results look like
In the response, you’ll find a buckets array under aggregations.distinct_trace_ids. Each bucket contains:
- The
key(your distincttrace_id) latest_timestamp.value(the newest timestamp for that trace_id)latest_record.hits.hits[0]._source(the full latest record withjob_nameand timestamp)
That's exactly the data you're after!
内容的提问来源于stack exchange,提问作者HShetty

