Elasticsearch双层嵌套对象聚合实现方法咨询
Got it, let's break down how to handle double nested aggregations with your sample data. First, a critical pre-requisite: your index mapping must explicitly mark both cat_a and cat_a.entry as nested types. Elasticsearch flattens object arrays by default, so without this, your aggregations will mix up values across nested objects incorrectly. Here's a valid mapping snippet for your data structure:
{ "mappings": { "properties": { "cat_a": { "type": "nested", "properties": { "position": { "type": "keyword" }, "tools": { "type": "keyword" }, "entry": { "type": "nested", "properties": { "tx_a": { "type": "keyword" }, "rx_a": { "type": "keyword" }, "number": { "type": "float" } } }, "basic": { "type": "boolean" } } } } } }
Once your mapping is set up correctly, you can build a chained nested aggregation to drill into the double-nested entry objects. Let's say you want to group by tx_a, then by each value in rx_a, and sum the number field for each combination. Here's the full aggregation DSL:
{ "size": 0, "aggs": { "sample_agg": { "nested": { "path": "cat_a" }, "aggs": { "inner_entry_level": { "nested": { "path": "cat_a.entry" }, "aggs": { "group_by_tx_a": { "terms": { "field": "cat_a.entry.tx_a" }, "aggs": { "group_by_rx_a": { "terms": { "field": "cat_a.entry.rx_a" }, "aggs": { "total_number": { "sum": { "field": "cat_a.entry.number" } } } } } } } } } } } }
Let's walk through each step:
- Top-level nested aggregation:
sample_aggtargets the first nested layer (cat_a) to access its inner fields. - Second nested aggregation:
inner_entry_leveldrills into the second nested layer (cat_a.entry)—this ensures we isolate eachentryobject under its parentcat_aitem, avoiding cross-object value mixing. - Term aggregation on
tx_a: Groups results by each uniquetx_avalue in the nestedentryobjects. - Term aggregation on
rx_a: Further splits eachtx_abucket by individual values in therx_aarray. - Sum aggregation on
number: Calculates the total of thenumberfield for everytx_a+rx_apair.
Example Aggregation Result
Based on your sample data, the response will match the structure you're expecting, looking something like this:
{ "aggregations": { "sample_agg": { "doc_count": 1, "inner_entry_level": { "doc_count": 2, "group_by_tx_a": { "buckets": [ { "key": "inside", "doc_count": 1, "group_by_rx_a": { "buckets": [ { "key": "soft_1", "doc_count": 1, "total_number": { "value": 0.018 } }, { "key": "soft_2", "doc_count": 1, "total_number": { "value": 0.018 } }, { "key": "soft_3", "doc_count": 1, "total_number": { "value": 0.018 } }, { "key": "soft_4", "doc_count": 1, "total_number": { "value": 0.018 } } ] } }, { "key": "out", "doc_count": 1, "group_by_rx_a": { "buckets": [ { "key": "soft_1", "doc_count": 1, "total_number": { "value": 0.0001 } }, { "key": "soft_3", "doc_count": 1, "total_number": { "value": 0.0001 } }, { "key": "soft_5", "doc_count": 1, "total_number": { "value": 0.0001 } }, { "key": "soft_7", "doc_count": 1, "total_number": { "value": 0.0001 } } ] } } ] } } } } }
The core idea here is chaining nested aggregations for each level of your nested object hierarchy—each nested aggregation "steps into" the next layer, ensuring your calculations are applied to the correct, isolated nested items.
内容的提问来源于stack exchange,提问作者meh

