如何基于关联文档实现Elasticsearch中人员的相关性搜索?
可行方案:基于Elasticsearch关联搜索人员并按报告相关性排序
一、可行性确认
完全可行,Elasticsearch提供了多种关联数据的查询能力,无需将报告内容冗余存储到Person索引中,可通过父子文档关系或跨索引查询+聚合排序两种核心方式实现需求。
二、具体实现方案
方案1:使用Elasticsearch Join数据类型(父子关系)
1. 数据建模
为Reports索引设置Join字段,指定person为父文档类型,report为子文档类型:
PUT /people_reports { "mappings": { "properties": { "person_id": {"type": "keyword"}, "content": {"type": "text"}, "join_field": { "type": "join", "relations": { "person": "report" } } } } }
将Person数据作为父文档写入,Reports作为子文档关联到对应Person。
2. 查询实现
使用has_child查询匹配包含目标关键词的报告,结合score_mode根据子文档的匹配度(如关键词出现频次、报告匹配数量)计算Person的最终得分:
GET /people_reports/_search { "query": { "has_child": { "type": "report", "query": { "match": { "content": "art director" } }, "score_mode": "sum" // 累加所有匹配报告的得分,也可选择max/avg等模式 } } }
3. Searchkick适配实现
在Ruby的Person模型中配置Searchkick,通过custom_query注入DSL逻辑:
class Person < ApplicationRecord searchkick index_name: "people_reports", mappings: { properties: { join_field: {type: "join", relations: {person: "report"}} # 配置Person自身字段 } } def search_data { # 填充Person自身字段 join_field: {name: "person"} } end # 自定义搜索方法 def self.search_by_report_keyword(keyword) search( "", custom_query: lambda do |body| body[:query] = { has_child: { type: "report", query: {match: {content: keyword}}, score_mode: "sum" } } end ) end end
方案2:跨索引查询+聚合排序(无需修改现有索引)
如果不想调整现有索引结构,可通过两步查询实现:
- 先搜索Reports索引,聚合统计每个Person ID对应的匹配得分总和:
GET /reports/_search { "size": 0, "query": { "match": { "content": "art director" } }, "aggs": { "person_scores": { "terms": { "field": "person_id", "size": 10000, "order": {"total_score": "desc"} }, "aggs": { "total_score": { "sum": { "script": "_score" } } } } } }
- 根据聚合结果查询People索引,并按预计算的得分排序:
def self.search_by_report_keyword(keyword) # 第一步:查询Reports并聚合得分 report_response = Elasticsearch::Client.new.search( index: "reports", body: { size: 0, query: {match: {content: keyword}}, aggs: { person_scores: { terms: {field: "person_id", size: 10000, order: {total_score: "desc"}}, aggs: {total_score: {sum: {script: "_score"}}} } } } ) # 提取Person ID与对应得分 person_scores = report_response.dig("aggregations", "person_scores", "buckets").map do |bucket| {id: bucket["key"], score: bucket["total_score"]["value"]} end # 第二步:查询People并按聚合得分排序 person_ids = person_scores.pluck(:id) score_map = Hash[person_scores.map { |ps| [ps[:id], ps[:score]] }] Person.search( "*", where: {id: person_ids}, custom_query: lambda do |body| body[:sort] = [ {_script: { type: "number", script: "params.scores.get(doc['id'].value)", params: {scores: score_map}, order: "desc" }} ] end ) end
三、方案对比
- 父子文档方案:查询效率更高,适合实时搜索场景,但需要调整索引结构,写入时需维护父子关联关系。
- 跨索引聚合方案:无需修改现有索引,兼容性好,但多了一次聚合查询,延迟略高,适合非实时或低并发场景。
内容的提问来源于stack exchange,提问作者Dauncing
相关产品推荐
相关产品推荐

