如何在Elasticsearch 6.5中按keyword字段特定值聚合数据
Hey there! Let's work through how to aggregate your data by the status values tucked inside the message keyword field. Since message is stored as a keyword (treated as a single, unbroken string), we can't directly split it to pull out those status numbers—so we'll use a runtime script to extract them first, then run our aggregation.
核心方案:用Painless脚本提取status值并聚合
Elasticsearch 6.5 supports the Painless scripting language, which we can use to regex-match the status numbers from your message strings. Here's a complete search request that does exactly what you need:
GET my_index/_search { "size": 0, // 不返回具体文档,只获取聚合结果,节省资源 "query": { "match": { "message": "status" } }, "aggs": { "status_aggregation": { "terms": { "script": { "source": """ // 用正则匹配message中的status数值 def statusMatcher = /status: (\d+)/.matcher(doc['message'].value); if (statusMatcher.find()) { // 返回匹配到的数字字符串 return statusMatcher.group(1); } // 处理未匹配到status的情况(可选) return 'unknown_status'; """, "lang": "painless" }, "size": 100 // 因为你有100余种状态,设置足够大的数值返回所有桶 } } } }
关键部分解释:
size: 0: We don't need to see individual hit documents, so setting this to 0 cuts down on unnecessary data transfer and speeds up the request.- Painless Script: The regex
/status: (\d+)/hunts for patterns likestatus: 123in themessagefield. When a match is found, we pull out the numeric part withgroup(1). If no match exists, we return a fallback value (unknown_status) to avoid dropping any documents. size: 100: By default, thetermsaggregation only returns the top 10 buckets. Since you have over 100 unique status values, we set this to a number large enough to capture all of them.
预期聚合结果:
You'll get a response structure just like what you wanted, looking something like this:
{ "aggregations" : { "status_aggregation" : { "buckets" : [ { "key" : "123", "doc_count" : 2 }, { "key" : "456", "doc_count" : 1 }, // 其他状态值的桶会依次列出 { "key" : "unknown_status", "doc_count" : 0 } ] } } }
优化建议(若可修改索引结构):
If this aggregation is something you'll run often, it's better to pre-extract the status value into a dedicated field at index time. This avoids runtime script overhead and makes aggregations faster. You can use an ingest pipeline with a grok processor to parse the message field and store the status as a separate keyword or integer field. For example:
PUT _ingest/pipeline/status_extractor { "processors": [ { "grok": { "field": "message", "patterns": ["status: %{NUMBER:status:int}"] } } ] }
Then when indexing documents, use this pipeline to automatically populate the status field. After that, your aggregation becomes as simple as:
GET my_index/_search { "size": 0, "query": { "exists": { "field": "status" } }, "aggs": { "status_counts": { "terms": { "field": "status", "size": 100 } } } }
内容的提问来源于stack exchange,提问作者Chen-Hao Ku

