如何在Elasticsearch function_score脚本中统计搜索词出现次数?
优化Elasticsearch多标签字段的匹配次数评分
原脚本性能差的原因
你当前的Painless脚本慢,核心原因是无法利用Elasticsearch的倒排索引优化:
- 脚本需要对每个匹配的文档,遍历所有搜索词,再逐个检查三个字段的标签值,时间复杂度为O(搜索词数量 × 三个字段的总标签数)
doc['字段名'].contains()是线性扫描该字段的所有值,没有利用索引的快速查找能力- 这种逐文档的计算逻辑完全在查询阶段执行,无法提前做预计算,数据量越大,性能越差
方案1:去重的匹配次数评分(每个搜索词出现就算1次,不管重复)
这种场景不需要脚本,直接用function_score结合filter和weight实现,完全利用倒排索引的快速匹配能力:
{ "query": { "function_score": { "query": { /* 你的基础查询条件 */ }, "functions": [ // 每个搜索词对应一个filter,匹配到任意字段就加1分 { "filter": { "bool": { "should": [ {"term": {"tags": "Tag1"}}, {"term": {"special_tags": "Tag1"}}, {"term": {"reserved_tags": "Tag1"}} ], "minimum_should_match": 1 } }, "weight": 1 }, { "filter": { "bool": { "should": [ {"term": {"tags": "Tag4"}}, {"term": {"special_tags": "Tag4"}}, {"term": {"reserved_tags": "Tag4"}} ], "minimum_should_match": 1 } }, "weight": 1 }, { "filter": { "bool": { "should": [ {"term": {"tags": "SpecialTag3"}}, {"term": {"special_tags": "SpecialTag3"}}, {"term": {"reserved_tags": "SpecialTag3"}} ], "minimum_should_match": 1 } }, "weight": 1 }, // 这里添加你的其他评分函数 ], "score_mode": "sum" } } }
优势
- 完全基于倒排索引的快速匹配,性能远高于脚本
- 每个搜索词只要在三个字段中任意一个出现,就累加1分,最终总分就是去重后的匹配次数
- 可以和其他
function_score函数无缝结合
方案2:包含重复的匹配次数评分(统计所有出现次数,含重复)
这种场景需要统计标签的实际出现次数,我们可以通过索引阶段合并字段+优化后的脚本来提升性能:
步骤1:修改索引映射,添加合并字段
用copy_to把三个标签字段的内容合并到一个新字段all_tags,索引时完成,不影响原字段的使用:
{ "mappings": { "properties": { "tags": { "type": "keyword", "copy_to": "all_tags" }, "special_tags": { "type": "keyword", "copy_to": "all_tags" }, "reserved_tags": { "type": "keyword", "copy_to": "all_tags" }, "all_tags": { "type": "keyword" }, // 你的其他字段 } } }
步骤2:优化后的Painless脚本
利用合并后的all_tags字段,结合Set的快速查找优化:
{ "query": { "function_score": { "query": { /* 你的基础查询条件 */ }, "functions": [ { "script_score": { "script": { "source": """ def matched = 0; def searchedTagSet = params.searchedTags.toSet(); def allTags = doc['all_tags'].values; for (tag : allTags) { if (searchedTagSet.contains(tag)) { matched++; } } return matched; """, "params": { "searchedTags": ["Tag1", "Tag4", "SpecialTag3"] } } } }, // 这里添加你的其他评分函数 ], "score_mode": "sum" } } }
优势
- 合并字段后,脚本只需要遍历一次所有标签,减少了多次字段访问的开销
- 把搜索词转成Set,
contains操作从O(n)变成O(1),提升遍历效率 - 相比原脚本,性能提升显著,同时保留了统计重复次数的能力
内容的提问来源于stack exchange,提问作者Nikhil Jagtap
相关产品推荐
相关产品推荐

