如何获取ElasticSearch中高亮字段的单独评分?
ElasticSearch 获取字段级评分解决方案需求
需求说明
现有查询可在索引的4个字段(title、description、subtitle、chapter)中搜索并返回高亮结果,但需要获取每个文档中高亮字段的单独评分,而非整个文档的评分。
业务场景:前端固定展示title字段,需从description、subtitle、chapter三个字段中筛选出与搜索关键词最相关的字段进行展示,因此依赖字段级评分判断相关性。
已尝试无效方案
- 使用
explain=true,结果过于复杂,无法便捷提取字段评分 - 尝试
sub_searches,未生效
当前查询代码
$filesQuery['body'] = [ 'size' => 5, 'query' => [ 'function_score' => [ 'query' => [ 'bool' => [ 'must' => [ [ 'bool' => [ 'should' => [ [ 'match' => [ 'title' => [ 'query' => $request['query'], 'boost' => 4 ], ], ], [ 'wildcard' => [ 'title' => [ 'value' => '*' . $request['query'] . '*', ], ], ], [ 'match' => [ 'description' => [ 'query' => $request['query'], 'boost' => 3 ], ], ], [ 'wildcard' => [ 'description' => [ 'value' => '*' . $request['query'] . '*', ], ], ], [ 'match' => [ 'chapters_v1.chapter' => [ 'query' => $request['query'], 'boost' => 3 ], ], ], [ 'wildcard' => [ 'chapters_v1.chapter' => [ 'value' => '*' . $request['query'] . '*', ], ], ], [ 'match' => [ 'subtitle' => [ 'query' => $request['query'], 'boost' => 1 ], ], ], [ 'wildcard' => [ 'subtitle' => [ 'value' => '*' . $request['query'] . '*', ], ], ], ], ], ], [ 'term' => [ 'team_id' => $teamId, ], ], ], ], ], ], ], 'collapse' => [ 'field' => 'file_id', ], 'aggs' => [ 'total_files' => [ 'cardinality' => [ 'field' => 'file_id', ], ], ], 'highlight' => [ 'fields' => [ 'title' => [ 'type' => 'plain', "fragment_size" => 70, "no_match_size" => 70, ], 'description' => [ 'type' => 'plain', "fragment_size" => 55, "no_match_size" => 55, ], 'subtitle' => [ 'type' => 'plain', 'number_of_fragments' => 1, 'fragment_size' => 30, ], 'chapters_v1.chapter' => [ 'type' => 'plain', 'number_of_fragments' => 1, 'fragment_size' => 30, ], ], ], 'track_total_hits' => true, ];
可行解决方案
方案思路:命名查询 + 脚本提取字段得分
通过给每个字段的查询(match/wildcard)添加唯一命名,再通过script_fields解析查询的_explanation,提取每个字段对应的查询得分总和,得到字段级评分。
修改后的查询代码
$filesQuery['body'] = [ 'size' => 5, 'explain' => true, // 必须开启,才能访问_explanation 'query' => [ 'function_score' => [ 'query' => [ 'bool' => [ 'must' => [ [ 'bool' => [ 'should' => [ // 给title的两个查询命名 [ 'match' => [ 'title' => [ 'query' => $request['query'], 'boost' => 4, '_name' => 'title_match' ], ], ], [ 'wildcard' => [ 'title' => [ 'value' => '*' . $request['query'] . '*', '_name' => 'title_wildcard' ], ], ], // 给description的两个查询命名 [ 'match' => [ 'description' => [ 'query' => $request['query'], 'boost' => 3, '_name' => 'description_match' ], ], ], [ 'wildcard' => [ 'description' => [ 'value' => '*' . $request['query'] . '*', '_name' => 'description_wildcard' ], ], ], // 给chapter的两个查询命名 [ 'match' => [ 'chapters_v1.chapter' => [ 'query' => $request['query'], 'boost' => 3, '_name' => 'chapter_match' ], ], ], [ 'wildcard' => [ 'chapters_v1.chapter' => [ 'value' => '*' . $request['query'] . '*', '_name' => 'chapter_wildcard' ], ], ], // 给subtitle的两个查询命名 [ 'match' => [ 'subtitle' => [ 'query' => $request['query'], 'boost' => 1, '_name' => 'subtitle_match' ], ], ], [ 'wildcard' => [ 'subtitle' => [ 'value' => '*' . $request['query'] . '*', '_name' => 'subtitle_wildcard' ], ], ], ], ], ], [ 'term' => [ 'team_id' => $teamId, ], ], ], ], ], ], ], 'collapse' => [ 'field' => 'file_id', ], 'aggs' => [ 'total_files' => [ 'cardinality' => [ 'field' => 'file_id', ], ], ], 'highlight' => [ 'fields' => [ 'title' => [ 'type' => 'plain', "fragment_size" => 70, "no_match_size" => 70, ], 'description' => [ 'type' => 'plain', "fragment_size" => 55, "no_match_size" => 55, ], 'subtitle' => [ 'type' => 'plain', 'number_of_fragments' => 1, 'fragment_size' => 30, ], 'chapters_v1.chapter' => [ 'type' => 'plain', 'number_of_fragments' => 1, 'fragment_size' => 30, ], ], ], // 新增script_fields,提取每个字段的得分 'script_fields' => [ 'title_score' => [ 'script' => [ 'source' => ' def score = 0; for (def exp : _explanation.details) { if (exp.description.contains("title_match") || exp.description.contains("title_wildcard")) { score += exp.value; } } return score; ' ] ], 'description_score' => [ 'script' => [ 'source' => ' def score = 0; for (def exp : _explanation.details) { if (exp.description.contains("description_match") || exp.description.contains("description_wildcard")) { score += exp.value; } } return score; ' ] ], 'chapter_score' => [ 'script' => [ 'source' => ' def score = 0; for (def exp : _explanation.details) { if (exp.description.contains("chapter_match") || exp.description.contains("chapter_wildcard")) { score += exp.value; } } return score; ' ] ], 'subtitle_score' => [ 'script' => [ 'source' => ' def score = 0; for (def exp : _explanation.details) { if (exp.description.contains("subtitle_match") || exp.description.contains("subtitle_wildcard")) { score += exp.value; } } return score; ' ] ], ], 'track_total_hits' => true, ];
方案说明
- 命名查询:给每个字段的
match和wildcard查询添加_name属性,方便后续在脚本中识别对应字段的查询得分。 - 开启explain:必须设置
'explain' => true,才能在脚本中访问查询的_explanation结构,从中提取各子查询的得分。 - script_fields提取得分:通过Painless脚本遍历
_explanation.details,累加对应字段的所有查询(match+wildcard)的得分,得到该字段的总评分。 - 结果使用:返回的每个文档会包含
fields字段,其中title_score、description_score等就是对应字段的单独评分,前端可根据这些值筛选最相关的字段展示。
优化建议
- 如果
wildcard查询的权重需要调整,可以在查询中添加boost属性,脚本会自动累加带boost的得分。 - 如果不需要
_explanation返回结果,可以在获取数据后忽略该字段,仅保留script_fields中的评分。
内容的提问来源于stack exchange,提问作者kishan maharana
相关产品推荐
相关产品推荐

