如何在Elasticsearch中聚合checker_conclusions对象的所有布尔字段
问题描述
示例文档结构:
{ "reference": "0001", "order_outlet": "Brighton", "order_amount": 1000, "status": "confirmed", "checker_conclusions": { "conclusion_1": true, "conclusion_2": false, "conclusion_3": false } }
需求:统计checker_conclusions对象下所有字段的true和false数量。目前仅能单独对单个字段做聚合,示例如下:
"aggs": { "checker_conclusions": { "terms": { "field": "checker_conclusions.is_pep" } } }
尝试过使用通配符"field":"checker_conclusions.*",以及通过dynamic_templates将所有字段映射为keyword类型,但均未成功。
期望得到的聚合结果格式:
{ "test_conclusion": { "doc_count_error_upper_bound": 0, "sum_other_doc_count": 0, "buckets": [ { "key": "conclusion_1", "buckets": [ { "key": "true", "doc_count": 1 }, { "key": "false", "doc_count": 3 } ] }, { "key": "conclusion_2", "buckets": [ { "key": "true", "doc_count": 4 }, { "key": "false", "doc_count": 2 } ] } ] }
解决方案
由于checker_conclusions是普通object类型,动态字段无法直接通过通配符实现批量聚合,提供两种可行方案:
方案1:脚本聚合(无需修改数据结构)
通过Painless脚本遍历checker_conclusions内的所有字段,提取字段名与对应布尔值后分层聚合:
{ "size": 0, "aggs": { "test_conclusion": { "terms": { "script": { "source": """ def entries = []; for (entry in params._source.checker_conclusions.entrySet()) { entries.add(entry.getKey() + '|' + entry.getValue()); } return entries; """, "lang": "painless" }, "size": 1000 }, "aggs": { "aggregated_results": { "scripted_metric": { "init_script": "state.buckets = [:];", "map_script": """ def parts = params._value.split('\\|'); def fieldName = parts[0]; def value = parts[1]; if (!state.buckets.containsKey(fieldName)) { state.buckets[fieldName] = ['true': 0, 'false': 0]; } state.buckets[fieldName][value] += 1; """, "combine_script": "return state.buckets;", "reduce_script": """ def finalBuckets = [:]; for (bucket in states) { for (fieldEntry in bucket.entrySet()) { def fieldName = fieldEntry.getKey(); def counts = fieldEntry.getValue(); if (!finalBuckets.containsKey(fieldName)) { finalBuckets[fieldName] = ['true': 0, 'false': 0]; } finalBuckets[fieldName]['true'] += counts['true']; finalBuckets[fieldName]['false'] += counts['false']; } } def result = []; for (entry in finalBuckets.entrySet()) { def subBuckets = [ {'key': 'true', 'doc_count': entry.getValue()['true']}, {'key': 'false', 'doc_count': entry.getValue()['false']} ]; result.add({'key': entry.getKey(), 'buckets': subBuckets}); } return result; """ } } } } } }
方案2:调整数据为嵌套类型(推荐,性能更优)
如果可以控制数据写入逻辑,将checker_conclusions改为嵌套数组结构,示例文档修改为:
{ "reference": "0001", "order_outlet": "Brighton", "order_amount": 1000, "status": "confirmed", "checker_conclusions": [ {"name": "conclusion_1", "value": true}, {"name": "conclusion_2", "value": false}, {"name": "conclusion_3", "value": false} ] }
同时在索引映射中声明checker_conclusions为nested类型:
{ "mappings": { "properties": { "checker_conclusions": { "type": "nested", "properties": { "name": {"type": "keyword"}, "value": {"type": "boolean"} } } } } }
之后即可通过嵌套聚合实现需求,逻辑更简洁,性能更稳定:
{ "size": 0, "aggs": { "test_conclusion": { "nested": { "path": "checker_conclusions" }, "aggs": { "by_conclusion_name": { "terms": { "field": "checker_conclusions.name.keyword", "size": 1000 }, "aggs": { "by_boolean_value": { "terms": { "field": "checker_conclusions.value" } } } } } } } }
内容的提问来源于stack exchange,提问作者Ray Stanton
相关产品推荐
相关产品推荐

