如何基于stats文档近4周访问量总和排序book搜索结果?
问题解答
能不能实现?
这事儿完全能实现,但你之前的查询逻辑有问题:你在整个索引里匹配title字段,但stats文档根本没有title,结果查询出来的只有book文档,可聚合又要用到只有stats才有的book_id字段,两者对不上,自然出不来正确结果。
要不要调整文档结构?
没必要用嵌套文档,更适合用**Elasticsearch的父子文档(Join类型)**来关联book和stats:
- 建索引的时候加个
join字段,指定book是父文档,stats是子文档,用book_id对应父文档的id就行。 - 这种结构既能保留两种文档的独立性,查父文档的时候还能高效关联子文档的统计数据。
要是你的数据更新不频繁,也可以直接把统计数据预存在book文档里:比如加个recent_4w_visits字段,定时计算近4周的访问总和更新进去,这样查询排序会更简单高效。
修正后的查询方法
方法一:基于父子文档结构(推荐)
假设已经把索引改成父子文档结构,父文档类型是book,子文档是stats,用下面的查询就行:
{ "size": 10, "query": { "bool": { "must": [ { "match": { "title": "my term" } } ] } }, "aggs": { "books_with_stats": { "terms": { "field": "id", // book文档的id字段 "size": 10 }, "aggs": { "recent_visits": { "children": { "type": "stats" // 子文档类型 }, "aggs": { "last_4_weeks": { "filter": { "range": { "visits_date": { "gte": "now-4w/w" } } }, "aggs": { "total_visits": { "sum": { "field": "visits_count" } } } } } } } } }, "sort": [ { "total_visits": { "order": "desc", "mode": "max" } } ] }
方法二:不调整文档结构的兼容方案
要是不想改结构,也可以用脚本把book的id和stats的book_id统一分组,再筛选计算:
{ "size": 0, "query": { "bool": { "should": [ { "match": { "title": "my term" } }, { "exists": { "field": "book_id" } } ] } }, "aggs": { "group_by_book": { "terms": { "script": { "source": "doc.containsKey('id') ? doc['id'].value : doc['book_id'].value" }, "size": 10 }, "aggs": { "has_book_match": { "filter": { "match": { "title": "my term" } } }, "recent_visits_sum": { "filter": { "bool": { "must": [ { "exists": { "field": "book_id" } }, { "range": { "visits_date": { "gte": "now-4w/w" } } } ] } }, "aggs": { "total_visits": { "sum": { "field": "visits_count" } } } }, "only_matching_books": { "bucket_selector": { "buckets_path": { "bookMatchCount": "has_book_match._count" }, "script": "params.bookMatchCount > 0" } } } } } }
这个查询会把book和对应的stats归到同一组,只保留符合标题搜索的book,再算出它们近4周的访问总和。
内容的提问来源于stack exchange,提问作者samuel
相关产品推荐
相关产品推荐

