Elasticsearch中基于Terms lookup文档属性实现过滤排序与聚合
好的,咱们来一步步拆解你的问题:你已经通过Terms lookup成功获取了用户集合关联的thing文档,但现在想基于collection里的time_added、condition这类属性,对这些关联文档做过滤、排序甚至聚合,还好奇这类似SQL左连接的需求,Elasticsearch到底适不适合。下面给你详细说说可行的方案和注意事项:
一、当然可以实现!三种方案供你选择
1. 优先推荐:数据冗余(Denormalization)——Elasticsearch的最佳实践
Elasticsearch天生适合扁平化的文档结构,最省心高效的方式就是把collection里的关联属性(比如condition、time_added)直接同步到对应的thing文档里。比如修改thing的结构:
{ "index":{ "_id": 1 } } { "title":"One thing", "condition": "fair", "time_added": "2017-08-07T09:07:15.000Z"} { "index":{ "_id": 3 } } { "title":"Three things", "condition": "good", "time_added": "2019-08-07T09:07:15.000Z"}
这样之后,过滤、排序、聚合就跟普通查询一样简单:
GET /index/thing/_search { "query": { "term": { "condition": "good" } }, "sort": [{"time_added": "desc"}], "aggs": { "condition_stats": { "terms": { "field": "condition" } } } }
这种方式性能拉满,完全没有跨文档查询的开销,是Elasticsearch处理关联场景的首选方案。
2. 不改结构:用Inner Hits + Nested查询反向关联
如果暂时没法修改thing的结构,可以反过来查询collection文档,利用inner_hits获取关联的thing,同时对items的属性做过滤和排序。比如筛选condition: good的关联文档,并按time_added倒序:
GET /index/collection/_search { "query": { "nested": { "path": "items", "query": { "term": { "items.condition": "good" } }, "inner_hits": { "query": { "terms": { "_id": { "index": "index", "type": "thing", "id": "{{_id}}", "path": "items.id" } } }, "sort": [{"items.time_added": "desc"}] } } } }
返回结果里的inner_hits字段就是符合条件的thing文档,同时已经按time_added排好序了。
3. 脚本实现(不推荐,仅适合小规模数据)
如果实在不想动结构,也可以用脚本在查询thing时,动态拉取collection里的关联属性,但这种方式性能很差——每个匹配的thing都要额外查询一次collection文档,只适合测试或者数据量极小的场景。
比如按time_added排序并过滤condition: good:
GET /index/thing/_search { "query": { "bool": { "filter": [ { "terms": { "_id": { "index": "index", "type": "collection", "id": 1, "path": "items.id" } } }, { "script": { "script": { "source": """ def collection = ctx._index.get('collection', params.col_id); def item = collection.items.stream().filter(it -> it.id == ctx._id).findFirst().orElse(null); return item != null && item.condition == 'good'; """, "params": { "col_id": "1" } } } } ] } }, "sort": [ { "_script": { "type": "date", "script": { "source": """ def collection = ctx._index.get('collection', params.col_id); def item = collection.items.stream().filter(it -> it.id == ctx._id).findFirst().orElse(null); return item != null ? item.time_added : null; """, "params": { "col_id": "1" } }, "order": "desc" } } ] }
注意:使用脚本需要提前在Elasticsearch配置中开启相关权限,生产环境谨慎使用。
二、Elasticsearch适合这类类似SQL左连接的需求吗?
答案是分场景:
- 如果是小规模数据、偶尔的关联查询:上述方案完全可以满足需求;
- 如果是大规模数据、高频关联操作:Elasticsearch的表现会不如关系型数据库,因为它的索引结构不是为复杂关联优化的;
- 核心原则:Elasticsearch的优势在于全文搜索和实时分析,处理关联场景时,尽量遵循数据冗余的最佳实践,减少跨文档查询的开销。
总的来说,你的需求完全可以实现,优先选择数据冗余的方式,既能保证性能,又能简化查询逻辑。
内容的提问来源于stack exchange,提问作者ComeAlongBort

