Elasticsearch:如何通过geohex_grid或geohash_grid返回完整文档详情?
Elasticsearch Geohash聚合获取文档详情的方案
要在geohash_grid聚合中获取文档详情,或者仅获取doc_count=1的桶的文档,可以通过嵌套子聚合的方式实现,以下是具体方案:
方案1:所有桶返回文档详情(按需控制数量)
在tiles_in_bounds聚合下嵌套top_hits子聚合,就能让每个geohash桶返回指定数量的文档详情。如果只想在doc_count=1时返回完整文档,把top_hits的size设为1即可。
修改后的请求示例:
POST vt_data_warehouse/_search?size=0 { "aggregations": { "tiles_in_bounds": { "geohash_grid": { "field": "position", "precision": 6, "bounds": { "top_left": "POINT (104.731344 11.541166)", "bottom_right": "POINT (104.89368038286129 11.464655721322128)" } }, "aggregations": { "bucket_docs": { "top_hits": { "size": 1 // 每个桶最多返回1条文档,doc_count=1时正好返回唯一的那条 } } } } } }
返回结果里,每个桶会新增bucket_docs字段,包含对应文档的完整内容。
方案2:仅返回doc_count=1的桶及对应文档
如果只关心doc_count=1的桶,可以结合bucket_selector聚合先过滤掉不符合条件的桶,再嵌套top_hits获取文档:
POST vt_data_warehouse/_search?size=0 { "aggregations": { "tiles_in_bounds": { "geohash_grid": { "field": "position", "precision": 6, "bounds": { "top_left": "POINT (104.731344 11.541166)", "bottom_right": "POINT (104.89368038286129 11.464655721322128)" } }, "aggregations": { "bucket_docs": { "top_hits": { "size": 1 } }, "filter_single_doc": { "bucket_selector": { "buckets_path": { "docCount": "_count" }, "script": "params.docCount == 1" } } } } } }
这个请求会只保留doc_count=1的桶,并且每个桶里包含对应的完整文档。
注意事项
top_hits的size参数可以根据需求调整,但不要设置过大,否则会增加聚合的内存开销和响应时间。- 如果索引数据量极大,这种嵌套聚合的性能会受影响,建议结合分片优化或者限制聚合的桶数量(通过
geohash_grid的size参数)。
内容的提问来源于stack exchange,提问作者Dara
相关产品推荐
相关产品推荐

