You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何针对完整嵌套location对象而非关键词实现Bucket聚合?

Got it, so you’re trying to bucket your documents by the entire nested locations object instead of just the location name—this way you can work with both the name and coordinates together in each bucket, right? Here are a couple of solid approaches to make that happen in Elasticsearch:

Method 1: Use a Script to Serialize the Full Location Object

You can create a unique key for each complete location by concatenating its name and coordinates into a single string. Then use that string as the basis for your terms aggregation. We’ll also add a top_hits sub-aggregation to return the full location details for each bucket:

{
  "aggs": {
    "nested_locations": {
      "nested": {
        "path": "locations"
      },
      "aggs": {
        "unique_locations": {
          "terms": {
            "script": {
              "source": "doc['locations.name.keyword'].value + '|' + doc['locations.coordinates.lat'].value + '|' + doc['locations.coordinates.long'].value"
            },
            "size": 1000 // Adjust based on how many unique locations you expect
          },
          "aggs": {
            "full_location_details": {
              "top_hits": {
                "size": 1,
                "_source": {
                  "includes": ["locations.name", "locations.coordinates"]
                }
              }
            }
          }
        }
      }
    }
  }
}

Why this works:

  • The script combines the location name (as a keyword to avoid text analysis issues) and coordinates into a unique string. Each distinct location will generate its own bucket.
  • The top_hits sub-aggregation pulls back the full, original location object for each bucket, so you can access both the name and coordinates easily.

Method 2: Use Composite Aggregation for Multi-Field Bucketing

If you prefer not to use scripts, the composite aggregation lets you create buckets based on combinations of multiple fields. This is great for large datasets since it supports pagination, and it avoids string concatenation:

{
  "aggs": {
    "nested_locations": {
      "nested": {
        "path": "locations"
      },
      "aggs": {
        "unique_locations": {
          "composite": {
            "sources": [
              { "location_name": { "terms": { "field": "locations.name.keyword" } } },
              { "latitude": { "terms": { "field": "locations.coordinates.lat" } } },
              { "longitude": { "terms": { "field": "locations.coordinates.long" } } }
            ]
          },
          "aggs": {
            "document_count": {
              "value_count": {
                "field": "_id"
              }
            }
          }
        }
      }
    }
  }
}

Why this works:

  • The composite aggregation creates a unique bucket for every combination of name, latitude, and longitude. This directly maps to your full locations object.
  • You can add sub-aggregations like value_count to get the number of documents associated with each location.

Quick Note on Coordinate Precision

If your coordinates are stored as floating-point numbers, you might run into issues with tiny precision differences (e.g., 36.78 vs 36.7800001) creating separate buckets. To fix this, format the coordinates to a consistent number of decimal places in the script:

"script": {
  "source": "doc['locations.name.keyword'].value + '|' + String.format('%.2f', doc['locations.coordinates.lat'].value) + '|' + String.format('%.2f', doc['locations.coordinates.long'].value)"
}

This ensures that minor precision variations don’t break your bucket grouping.

内容的提问来源于stack exchange,提问作者Carasel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:05:29